Docs
Chat Completions, OpenAI Responses, and Anthropic Messages through one base URL. The router translates tool calls and tool results between the SDK formats.
Quickstart
export OPENAI_BASE_URL=/v1 export OPENAI_API_KEY=ir-live-… export OPENAI_MODEL=openai/gpt-6.1-sol-pro
curl /v1/chat/completions \
-H "authorization: Bearer ir-live-…" \
-H "content-type: application/json" \
-d '{"model":"openai/gpt-6.1-sol-pro","messages":[{"role":"user","content":"hello"}]}'
Get a key
No account, no email, no KYC. Get a key and save it with your dashboard link — both are secrets. Top up any amount in USDG (Robinhood Chain) or USDC (Base): connect a wallet on the dashboard and sign one message. Gasless — the facilitator relays.
The API
| Endpoint | What it does |
|---|---|
| GET /v1/models | The published catalog. Every row carries its price (provider cost + 5%, included) and the pricing law. |
| POST /v1/chat/completions | Chat completions, streaming and not. Tools, JSON mode, vision-content — whatever the model supports, in OpenAI shape. |
| POST /v1/responses | Responses requests and event streams, including function tools and custom tool calls. Uses the same billing and routing as Chat Completions. |
| POST /v1/messages | Anthropic Messages requests and event streams, including tool use and results. Supports bearer keys and the SDK's X-Api-Key header. |
| GET /quote?model=&chars=&maxTokens= | What a body would cost, before you send it. |
| GET /healthz | Liveness + rails. |
Response headers
X-IR-Model names the model that served you. X-IR-Route is the upstream. For non-streaming calls, X-IR-Billed-Usd reports the amount charged: measured token usage for keys (an estimate if usage is missing), or the prepaid quote for direct wallet calls.
Auto tiers
deadcat/auto selects a catalog default, deadcat/auto:cheap a budget default, and deadcat/auto:code a coding default. Each catalog row names its current target and rates. Tool support depends on that model.
Pay per call (agents)
Skip the key entirely: POST without payment and receive an x402 challenge, then retry with the signed X-PAYMENT header. The exact quote includes our 5% markup, uses estimated input tokens and the requested maximum output, and has a $0.001 floor. It settles before inference; unused output is not refunded. Use a topped-up key for measured token billing. Rails: USDG on Robinhood Chain, USDC on Base — EIP-3009; the facilitator pays gas.
curl -sS /v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"deepseek/deepseek-v4-flash","messages":[{"role":"user","content":"say hi"}],"max_tokens":32}'
# → 402 with accepts[] rows (USDG / USDC). Sign one, then:
curl -sS /v1/chat/completions \
-H 'content-type: application/json' \
-H "X-PAYMENT: <base64 {x402Version,scheme,network,payload}>" \
-d '{…same body…}'
Client setups
Set the OpenAI SDK base URL to /v1, paste your API key, and choose a published model ID. For Anthropic SDKs, use this site's origin as the base URL. Messages and Responses use the existing x402-tokens adapters; a non-streaming provider's completed reply is emitted as the requested SDK's stream events. Responses requests must include the conversation input; server-side previous_response_id storage is not implemented.
Batch variants are excluded from this synchronous API. They require asynchronous submission and result retrieval; the current upstream gateways expose no batch API. Use a model listed in the published catalog for chat, tools, and streaming.
Fair use
The free lane is rate-limited per IP for trying things. Paid calls require credit, and upstream authorizations have a $1 per-call ceiling. Gateway usage events record model, cost, route and payment metadata without message text. Providers have their own data policies.