{"openapi":"3.1.0","info":{"title":"x402 LLM Gateway","version":"0.1.0"},"paths":{"/":{"get":{"summary":"Service Card","description":"Free service description — the agent's 'storefront'.\n\nWritten for an AI agent deciding whether to call this, not a human shopper:\nwhat it is, why it's worth using (private + no account), how to pay (x402),\na copy-paste example per tier, and the free path when there's no wallet yet.","operationId":"service_card__get","responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}}}}},"/v1/models":{"get":{"summary":"List Models","description":"Free model discovery, proxied from LM Studio.","operationId":"list_models_v1_models_get","responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}}}}},"/v1/extended":{"post":{"summary":"Chat Completions","description":"Paid endpoint (3 tiers, one handler). Proxies to LM Studio and returns an\nOpenAI-shaped response. The caller's max_tokens is CLAMPED to the tier's\noutput cap (never rejected after payment); the x402 stamp records the tier,\nthe charge, and - when clamped - which higher tier serves the full length.","operationId":"chat_completions_v1_extended_post","responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}}}}},"/v1/long":{"post":{"summary":"Chat Completions","description":"Paid endpoint (3 tiers, one handler). Proxies to LM Studio and returns an\nOpenAI-shaped response. The caller's max_tokens is CLAMPED to the tier's\noutput cap (never rejected after payment); the x402 stamp records the tier,\nthe charge, and - when clamped - which higher tier serves the full length.","operationId":"chat_completions_v1_long_post","responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}}}}},"/v1/chat/completions":{"post":{"summary":"Chat Completions","description":"Paid endpoint (3 tiers, one handler). Proxies to LM Studio and returns an\nOpenAI-shaped response. The caller's max_tokens is CLAMPED to the tier's\noutput cap (never rejected after payment); the x402 stamp records the tier,\nthe charge, and - when clamped - which higher tier serves the full length.","operationId":"chat_completions_v1_chat_completions_post","responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}}}}},"/v1/embeddings":{"post":{"summary":"Embeddings","description":"Paid embedding endpoint (own Bazaar listing). Proxies to LM Studio's local\nnomic-embed encoder. OpenAI-compatible: {input: string | string[]} ->\n{data: [{embedding: [...]}]}. Same 'never reject after payment' rule as the\nchat tiers: oversized batches are CLAMPED to the first max_batch texts and\nthe x402 stamp says so.","operationId":"embeddings_v1_embeddings_post","responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}}}}},"/v1/vision":{"post":{"summary":"Vision","description":"Paid VISION endpoint (own Bazaar listing). Uses the same local VLM as the\nchat tiers (the loaded Qwen model is a vision-language model), so agents pay\nfor private image understanding. Accepts the OpenAI messages format (image_url\nparts) or a simple {image, prompt}. Images are CLAMPED to the first max_images\n(never rejected after payment); the x402 stamp records the tier, the charge and\nany clamp. Reuses _serve_inference for the thinking-guard + empty-retry.","operationId":"vision_v1_vision_post","responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}}}}},"/v1/trial":{"post":{"summary":"Trial Chat","description":"Free trial call - served only while the promo window is OPEN.\nSame model, same quality guarantees (clamp + thinking guard), smaller cap,\nrate-limited, and stamped with the paid tiers it previews.","operationId":"trial_chat_v1_trial_post","responses":{"200":{"description":"Successful Response","content":{"application/json":{"schema":{}}}}}}}}}