# RouterPlus: all documentation > One API key for many LLM providers. OpenAI-compatible /v1/chat/completions and > Anthropic-compatible /v1/messages on one gateway. Pass-through pricing, > per-request cost on the wire, honest failover that never switches mid-answer. Gateway base URL: https://api.routerplus.com Create an account in your browser: https://app.routerplus.com/signup. Complete secure verification, then create an API key at https://app.routerplus.com/console/keys. The index is https://app.routerplus.com/llms.txt. Every page follows, in sidebar order, each under its source URL. --- Source: https://app.routerplus.com/docs/overview.md # Overview RouterPlus is an LLM gateway: one API key and one prepaid balance in front of models from multiple providers. It speaks the two wire dialects your code already speaks — an OpenAI-compatible surface and an Anthropic-compatible surface — and every chat model in the catalog is callable from **either** one (image models answer on [POST /v1/images/generations](/docs/api-images), video models on [POST /v1/videos](/docs/api-videos)). The gateway translates requests, streams, and errors between dialects; your SDK never notices. Pricing is pass-through: you pay the listed per-token rates, and the exact cost of every billed request is written into the response itself. ## Two wire surfaces, one key | Surface | Base URL | Completion endpoint | Works with | |---|---|---|---| | OpenAI-compatible | `https://api.routerplus.com/v1` | `POST /v1/chat/completions` | OpenAI SDKs, Codex CLI, anything speaking the OpenAI wire format | | Anthropic-compatible | `https://api.routerplus.com` | `POST /v1/messages` | Anthropic SDKs, Claude Code | The same key authenticates on both, as `Authorization: Bearer` or `x-api-key` — see [Authentication](/docs/authentication). The surface does not constrain the model: OpenAI-format code can call Claude models, Anthropic-format code can call GPT models. There is no lock between the dialect you speak and the model you get. ## Your first call Open [Sign up](https://app.routerplus.com/signup) in your browser, complete Clerk authentication and email verification, and save the first API key shown after sign-in. Add paid credits in [Billing](https://app.routerplus.com/console/billing) before your first call; signup does not add free credit. Existing customers can [sign in](https://app.routerplus.com/login) and create another key at [API keys](https://app.routerplus.com/console/keys). We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match. ```bash # Export the key saved from the browser. export TM_API_KEY=tm_vk_... # Call a model — streamed, billed, cost on the wire. curl -N https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"claude-haiku-4-5","stream":true,"max_tokens":60,"messages":[{"role":"user","content":"hello"}]}' ``` The final usage chunk of the stream carries token counts and `cost` — the exact USD amount this request debited. When the balance is spent, visit [Billing](https://app.routerplus.com/console/billing) to add credits. Or point your existing SDK at the gateway — the only changes are the base URL and the key: ```python import os from openai import OpenAI client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) r = client.chat.completions.create( model="claude-haiku-4-5", # yes — a Claude model over the OpenAI wire format max_tokens=60, messages=[{"role": "user", "content": "hello"}], ) print(r.model_dump()["usage"]["cost"]) # exact USD debit for this request ``` ```python import os import anthropic client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"]) msg = client.messages.create( model="gpt-4o-mini", # and a GPT model over the Anthropic wire format max_tokens=60, messages=[{"role": "user", "content": "hello"}], ) ``` ## Beyond the API The site uses Clerk for browser signup and sign-in. Its pages run on the same catalog, routing, prices and ledger as the API: - [Playground](https://app.routerplus.com/playground) — chat with any catalog model, compare up to three side by side, judge with a decision model, and render images and video. See [Playground](/docs/playground). - [Optimize](https://app.routerplus.com/model-search) — search for a cheaper model or mix that scores as well as yours on your own Braintrust evals, then deploy it. See [Optimize](/docs/model-search). - [Endpoints](https://app.routerplus.com/endpoints) — your organization's [dedicated endpoints](/docs/dedicated-endpoints), then the models you deployed from Optimize, each a `tm/...` model id your keys can call. - [Console](https://app.routerplus.com/console) — usage, logs, API keys, provider connections, integrations, credits, spend caps and rate limits, one page per section. ## What makes this gateway different ### The exact cost of every request is on the wire Every billed response carries `usage.cost` (USD) — in the final usage chunk on streams, in the response body otherwise. It is computed with the same integer micro-USD math the ledger settles with, and it covers the **full** request debit, including attempts that failed over before your answer started. You can recompute your bill from the wire at any time. Money math rounds in your favor: reservations round up, settlement rounds down. `GET /v1/usage` lists your balance and recent requests with per-request cost; `GET /v1/generation?id=` audits every physical attempt behind one request. One documented limit: an Anthropic-surface stream that dies mid-answer has no legal wire slot for usage, so `/v1/generation` is the recomputation path there. ### Failover with a commit boundary Before any output has reached you, upstream failures fail over between deployments with zero backoff — invisibly, except for the `x-tm-attempts` header that counts physical dispatches. The moment the first real output reaches you, the request is **committed** to that provider forever. If the provider dies after that, you get one terminal error event inside the stream — never a silent restart on a different provider, never a mid-answer switch. A content-policy refusal is never rerouted to another provider, period: rerouting a refusal would be laundering it. During long silences (slow reasoning models), the stream carries keep-alive frames every 15 seconds — an SSE comment on the OpenAI surface, a native `ping` event on the Anthropic surface — so proxies don't kill a healthy stream. ### Strict where silence would cost you money - An unknown model id is an honest **404 naming the id you sent** — never a silent substitution with a model you didn't choose. - `n>1` is rejected with a 400 rather than billing you for a garbled single-choice response. - Content that cannot cross a dialect boundary (image and audio parts, in v1) is a typed 400 naming the field whenever translation is required — never silently dropped. So is a part the model does not take, such as an image to a model that reads text only. A parameter the other dialect has no mapping for (JSON mode on an Anthropic deployment, say) is a typed 400 too. Unknown top-level parameters are dropped on every route, and every drop is recorded: the `x-tm-dropped-params` header names them, and so does the request's audit row. - When every deployment serving a model is cooling down, you get a 503 with `retry-after` — not a permanent-looking 404. ### Content-free by design Prompts and completions are never stored by the gateway. Metering, the usage endpoints, and telemetry record metadata only: token counts, timings, outcomes, cost. The optional `HTTP-Referer` and `X-Title` headers identify your app for analytics without exposing request content. One exception, stated where it applies: [Optimize](/docs/model-search) stores the eval cases and candidate answers of a search, for your organization, until you delete the search. See [Data policy](/docs/data-policy). ### Uptime numbers that admit their sample size Each model's page at `https://app.routerplus.com/models/` shows per-provider uptime over the last 24 hours. For a provider we call directly, it shows a percentage only once that provider has at least 100 counted requests. Below that it shows `n<100`, because a percentage over a handful of requests is noise dressed as data. Buyer-caused failures (your 400s, your cancelled streams) never count against a provider's uptime. See [The uptime figure is allowed to say nothing](/docs/models#the-uptime-figure-is-allowed-to-say-nothing). ## Endpoints at a glance Gateway (`https://api.routerplus.com`) — all authenticated: | Endpoint | What it does | |---|---| | `POST /v1/chat/completions` | OpenAI-compatible completions, streaming and not | | `POST /v1/messages` | Anthropic-compatible messages, streaming and not | | `POST /v1/messages/count_tokens` | Anthropic token counting, unbilled; models with an Anthropic-dialect deployment only | | `GET /v1/models` | Models your key can call. OpenAI list shape by default; Anthropic shape when you send an `anthropic-version` header | | `POST /v1/images/generations` | OpenAI Images API: base64 in the body | | `POST /v1/videos` | OpenAI Videos API: a job, then `GET /v1/videos/{id}` and `GET /v1/videos/{id}/content` | | `GET /v1/usage` | Balance, credited/spent totals, recent requests with cost | | `GET /v1/generation?id=` | Per-attempt audit for one request: tokens, provenance, cost, outcome | | `POST /v1/route` | Explains how a request would route, without dispatching it. See [Routing policies](/docs/routing-policies) | | `GET /v1/limits?model=` | The effective rate limits for your key. See [Limits and capacity](/docs/admission) | Site (`https://app.routerplus.com`): | Endpoint | What it does | |---|---| | `/signup` | Browser signup through Clerk, including email verification | | `GET /api/models.json` | Public catalog: ids, prices, context windows, who serves each model. No auth | | `/login` | Browser sign-in through Clerk | | `/playground` | Chat, decision, images and video with any catalog model | | `/model-search` | Optimize: model search on your Braintrust evals | | `/endpoints` | Your dedicated endpoints and your deployed `tm/...` models | | `/console` | Overview; then `/console/usage`, `/console/logs`, `/console/keys`, `/console/byok`, `/console/integrations`, `/console/billing`, `/console/limits` | | `/llms.txt` | Machine-readable docs index for agents | | `/llms-full.txt` | Every docs page as raw markdown, in one file | ## Where next - [Quickstart](/docs/quickstart) — browser signup to a streamed API call in five steps. - [Agent integration guide](/docs/agent-integration) — one page an AI agent reads to build the gateway into your product, with the best practices. - [Authentication](/docs/authentication) — virtual keys, both header forms, browser signup, what a 401 looks like. - [Errors & remediation](/docs/errors) — every `error_type`, what it means, and exactly what to do, including retry discipline for agents. - [Wire compatibility & billing contract](/docs/compat) — the conventions your bill is computed from; changing that page requires a recorded decision. - [Migrate in one prompt](/docs/migration) — repoint the OpenAI SDK, Anthropic SDK, Claude Code, or Codex CLI with a single paste. - [Install (agent runbook)](/docs/install) — hand it to your coding agent; it configures your client after browser signup. > [!TIP] > If a coding agent is doing the work, point it at > [https://app.routerplus.com/docs/agent-integration.md](https://app.routerplus.com/docs/agent-integration.md) to build the > gateway into a product, or at [Install](/docs/install) to wire a coding tool. > [https://app.routerplus.com/llms.txt](https://app.routerplus.com/llms.txt) is the index it fetches > first; [https://app.routerplus.com/llms-full.txt](https://app.routerplus.com/llms-full.txt) holds > every page in one file. --- Source: https://app.routerplus.com/docs/quickstart.md # Quickstart Signup to a streamed, billed model call in five steps. You need a browser, `curl` and an email address. The marketplace is one API key and one prepaid balance in front of two wire surfaces: | Surface | Endpoint | Works with | |---|---|---| | OpenAI-compatible | `POST https://api.routerplus.com/v1/chat/completions` | OpenAI SDKs, anything OpenAI-shaped | | Anthropic-compatible | `POST https://api.routerplus.com/v1/messages` | Anthropic SDKs, Claude Code | Every chat model in the catalog is callable from **both** surfaces — the gateway translates requests, streams, and errors in either direction. Image models have their own route, [POST /v1/images/generations](/docs/api-images). Prices are pass-through, and every billed response carries `usage.cost` in USD, so you can recompute your bill from the wire. > [!TIP] > Setting up with a coding agent instead of by hand? Hand it the runbook on > [Coding agents](/docs/install) — complete browser signup first, then let the agent configure your client. ## 1. Sign up in the browser and save your key Open [Sign up](https://app.routerplus.com/signup) and complete Clerk authentication and email verification. Your first verified sign-in shows your first `tm_vk_` API key. Save it then: only its hash is stored, so the raw key cannot be shown again. Add paid credits in [Billing](https://app.routerplus.com/console/billing) before your first call. Signup does not add free credit. We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match. Already have an account? [Sign in](https://app.routerplus.com/login) with the same verified email and create a key at [API keys](https://app.routerplus.com/console/keys). Your existing balance and API keys are preserved. Export the key: ```bash export TM_API_KEY=tm_vk_... ``` Account creation requires the browser flow; the former `POST /v1/signup` route has been removed. Once you have a key, all gateway calls below work from your terminal or SDK. See [Authentication](/docs/authentication). ## 2. Pick a model ```bash # Public catalog — no auth: ids, prices, context windows, provider retention curl -s https://app.routerplus.com/api/models.json # Authenticated — what your key can call, in your SDK's native list shape curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/models ``` `GET /v1/models` returns the OpenAI list shape by default, and the Anthropic shape when you send an `anthropic-version` header. See [Models & catalog](/docs/models). > [!NOTE] > Model access is catalog-exact. An id that isn't listed returns a `404` with > `error_type: model_unavailable`, echoing the id you asked for. The gateway never > substitutes a "close enough" model — no silent aliasing, ever. If a request > fails on the model id, the fix is the id, not a hidden routing preference. ## 3. First streamed call — OpenAI surface A Claude model over the OpenAI wire format, to prove the translation is real: ```bash curl -N https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" \ -H "content-type: application/json" \ -d '{ "model": "claude-haiku-4-5", "stream": true, "max_tokens": 60, "messages": [{"role": "user", "content": "Say hello in five words."}] }' ``` The stream arrives in a fixed order: a role-priming delta, content deltas, a finish chunk, a **usage chunk**, then `data: [DONE]`. The usage chunk is your bill: ```json { "prompt_tokens": 13, "completion_tokens": 9, "total_tokens": 22, "prompt_tokens_details": { "cached_tokens": 0, "cache_write_tokens": 0 }, "completion_tokens_details": { "reasoning_tokens": 0 }, "cost": 0.000058 } ``` `cost` is USD, computed with the same integer micro-USD math the ledger settles with, and it covers the full request — including any attempts that failed over before your answer started. Recompute it from the token counts and the public prices any time; [Pricing & billing](/docs/pricing) has the exact math. Same call with the OpenAI Python SDK — the only changes from stock OpenAI are `base_url` and the key: ```python import os from openai import OpenAI client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) stream = client.chat.completions.create( model="claude-haiku-4-5", max_tokens=60, stream=True, messages=[{"role": "user", "content": "Say hello in five words."}], ) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="") if chunk.usage: # the final chunk before [DONE] print(f"\ncost: ${chunk.usage.cost}") ``` And TypeScript: ```typescript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY, }); const stream = await client.chat.completions.create({ model: "claude-haiku-4-5", max_tokens: 60, stream: true, messages: [{ role: "user", content: "Say hello in five words." }], }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); } ``` > [!TIP] > Before dispatch, the gateway reserves the worst-case cost of the call — > roughly (estimated input + `max_tokens`) at the model's prices — and settles > down to observed usage afterward. On a small trial balance, set a sane > `max_tokens` (it defaults to 4096, and 32,768 is the most a request may ask for): a huge value can make the reservation > exceed your balance and return `insufficient_quota` before any provider is > called. Details in [Rate limits & spend caps](/docs/limits). ## 4. Same key, Anthropic surface A GPT model over the Anthropic wire format — the translation runs both ways: ```bash curl -N https://api.routerplus.com/v1/messages \ -H "x-api-key: $TM_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{ "model": "gpt-4o-mini", "stream": true, "max_tokens": 60, "messages": [{"role": "user", "content": "Say hello in five words."}] }' ``` The final `message_delta` event carries the usage: `input_tokens`, `output_tokens`, `cache_read_input_tokens`, `cache_creation_input_tokens`, and the same `cost` field in USD. With the Anthropic Python SDK: ```python import os import anthropic client = anthropic.Anthropic( base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"], ) with client.messages.stream( model="gpt-4o-mini", max_tokens=60, messages=[{"role": "user", "content": "Say hello in five words."}], ) as stream: for text in stream.text_stream: print(text, end="") ``` > [!NOTE] > Both auth header styles work on both endpoints: `Authorization: Bearer` and > `x-api-key`. Use whichever your SDK sends — no per-surface key juggling. See > [Authentication](/docs/authentication). ## 5. What did it cost? ```bash curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usage ``` ```json { "balance_usd": 4.999942, "credited_usd": 5.0, "spent_usd": 0.000058, "recent_attempts": [ { "request_id": "5a2e…", "deployment": "anthropic", "model": "claude-haiku-4-5", "outcome": "completed", "usage_provenance": "observed", "input_tokens": 13, "output_tokens": 9, "cost_usd": 0.000058, "billing_source": "house", "at": "2026-09-04T10:14:03.201Z" } ] } ``` The example is trimmed: each attempt also carries its price snapshot and inference cost, described on [GET /v1/usage](/docs/api-usage). `recent_attempts` lists your last 20 physical attempts. For the full audit of one request — every attempt including failovers, cache splits, reserved vs settled cost — use the `x-request-id` header from any response: ```bash curl -s -H "Authorization: Bearer $TM_API_KEY" \ "https://api.routerplus.com/v1/generation?id=REQUEST_ID" ``` Both endpoints are metadata-only: token counts, timings, outcomes, cost. Prompts and responses are never stored — see [Data policy](/docs/data-policy) and [GET /v1/usage & /v1/generation](/docs/api-usage). ## Deliberate strictness you'll notice - Unknown model ids get a `404`, never a substitute ([Routing & failover](/docs/routing)). - `n > 1` is rejected with a `400 invalid_request` rather than billing you for a garbled single-choice response. - First keys shown after verified sign-in start at 20 requests/minute. Until your organization buys credit, all its keys together stay at 20 requests/minute; after the first purchase, keys created in the console get 300. Every organization also has token-per-minute and concurrency limits ([Rate limits & spend caps](/docs/limits)). - Every error body carries a stable `error_type`, on both surfaces, in your SDK's native error shape ([Errors](/docs/errors)). ## Next steps - [Agent integration guide](/docs/agent-integration) — the page an AI agent follows to build the gateway into your product - [Coding agents](/docs/install) — point Claude Code or Codex at the gateway, then let an agent configure your client after browser signup - [Migrate](/docs/migration) — repoint an existing codebase in one prompt - [Streaming](/docs/streaming) — event order, keep-alives, and mid-stream failure semantics - [Pricing & billing](/docs/pricing) — the exact math behind `usage.cost` - [Wire compatibility](/docs/compat) — the contract your bill is computed from --- Source: https://app.routerplus.com/docs/agent-integration.md # Agent integration guide This page is for an AI coding agent that builds RouterPlus into a product's code. It covers the whole integration in order: the key, the SDK, model calls, provider pinning, sessions and prompt caching, errors, limits and cost. Each step links to the reference page that has every detail. Every docs page is also raw markdown: add `.md` to its URL, for example `https://app.routerplus.com/docs/errors.md`. The index of all pages is `https://app.routerplus.com/llms.txt`. All pages in one file: `https://app.routerplus.com/llms-full.txt`. > [!NOTE] > **For agents.** Do not guess model ids, request fields or headers. Copy them from this page or from the live catalog. If an error message disagrees with your plan, trust the error message and read the page it names. Some steps need a person. They are listed in [Steps only a human can do](#steps-only-a-human-can-do). Stop and ask at those steps. ## The short version 1. Point the OpenAI SDK at `https://api.routerplus.com/v1`, or the Anthropic SDK at `https://api.routerplus.com`. Read the key from `TM_API_KEY`, on the server only. 2. Copy model ids exactly from the catalog. The gateway never substitutes a model: an unknown id is a 404. 3. Send `max_tokens` on every request, from 1 to 32,768. Without it, the gateway uses 4,096. 4. When you need provider features such as prompt caching, call Claude models on `/v1/messages` and other models on `/v1/chat/completions`. 5. Keep the start of the prompt identical from turn to turn, and only append. This keeps prompt caches warm. 6. Pin a provider only when you must, with the `provider` object. Check the plan first with `POST /v1/route`. 7. Switch on `error_type`, never on message text. Retry `rate_limit` after `retry-after`. Stop on `insufficient_quota` and tell a human. 8. There is no idempotency key. A retry is a new request, and it is billed again if it reaches a provider. 9. Keep at most 8 requests in flight per organization, unless you raised that limit. 10. Log `x-request-id` and `usage.cost` for every call. ## At a glance | Item | Value | |---|---| | OpenAI-compatible base URL | `https://api.routerplus.com/v1`, for `POST /v1/chat/completions` | | Anthropic-compatible base URL | `https://api.routerplus.com`, for `POST /v1/messages` (the SDK adds `/v1`) | | API key | `tm_vk_` followed by 48 hex characters | | Auth header | `Authorization: Bearer ` or `x-api-key: `, on every route | | Model list | `GET https://api.routerplus.com/v1/models` with your key, or `https://app.routerplus.com/api/models.json` (public, with prices) | | Cost of a call | `usage.cost`, in USD, in every billed response | | Audit of a call | the `x-request-id` response header, then `GET https://api.routerplus.com/v1/generation?id=` | | Images, video and decisions | `POST /v1/images/generations`, `POST /v1/videos` and `POST /v1/decisions` | ### What is not available Do not build on these. They do not exist on the gateway today: - The OpenAI Responses API, embeddings, audio, moderation, files and batches. These paths return 404 `not_found`. Use Chat Completions or Messages. - Server-side conversation memory and idempotency keys. (Session ids exist: see [Sessions and caching](#sessions-and-caching).) - More than one answer per request (`n` above 1 is a 400). - Model aliases and model fallback lists. OpenRouter's `models` array and `route` field are dropped. - Sorting providers by price or speed, and price limits. `provider.sort` and `provider.max_price` are a 400. - Provider keys or endpoints in a request. `api_key`, `base_url` and `connection_id` in the body are a 400. ## 1. Get a key and keep it safe | Key | Where it comes from | Requests per minute | Use it for | |---|---|---|---| | First trial key | Shown after verified browser sign-in through Clerk | 20 | The first test calls | | Console key | [https://app.routerplus.com/console/keys](https://app.routerplus.com/console/keys) | 300 | Production, staging and CI | | Identity API key | `POST https://app.routerplus.com/api/identity/keys` | 300 | One key for each end user or service (step 7) | To get a key, the user opens [Sign up](https://app.routerplus.com/signup), completes Clerk authentication and email verification, and saves the first key shown after sign-in. Add paid credits in [Billing](https://app.routerplus.com/console/billing); signup does not grant free credit. Existing customers sign in at [https://app.routerplus.com/login](https://app.routerplus.com/login) using the same verified email and create a key at [API keys](https://app.routerplus.com/console/keys). Each raw key is shown once. We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match. Browser signup is required; the former programmatic signup and local magic-link submission routes have been removed. Once `TM_API_KEY` is available, the agent can configure clients and make API calls normally. Rules for the key: 1. Keep the key on the server. Read it from the `TM_API_KEY` environment variable or from your secret store. Never commit it, log it or put it in a URL. 2. Do not put the key in a browser or a mobile app. Anyone can copy it from there. The gateway accepts cross-origin requests, so a copied key works from any web page. If a browser-only internal tool must call the gateway, give it its own key with a low monthly spend cap. 3. Use one key for each service and each environment. Spend, caps and logs are kept per key, and you can disable one key without stopping the others. 4. Use a console key in production. A trial key allows only 20 requests per minute. 5. To rotate a key: create the new key, deploy it, then disable the old key in the console. A disabled key stops working within about 5 seconds. The console cannot enable it again. An owner or admin can, with the identity API. Details: [Authentication](authentication.md). ## 2. Connect the SDK Change two values in the client you already use: the base URL and the key. Set both in code. Do not depend on `OPENAI_BASE_URL` or `ANTHROPIC_BASE_URL` in the environment. When that variable is missing, the SDK sends the request, with its key, to the provider's own API instead of the gateway. Python — the OpenAI SDK: ```python import os from openai import OpenAI client = OpenAI( base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"], # Optional: these name your app in usage analytics. They carry no content. default_headers={"HTTP-Referer": "https://your-app.example", "X-Title": "Your App"}, ) ``` TypeScript — the OpenAI SDK: ```typescript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY, defaultHeaders: { "HTTP-Referer": "https://your-app.example", "X-Title": "Your App" }, }); ``` With the Anthropic SDK, the base URL has no `/v1`: ```python import os import anthropic client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"]) ``` ```typescript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ baseURL: "https://api.routerplus.com", apiKey: process.env.TM_API_KEY }); ``` ### Which surface to call Every chat model answers on both surfaces, because the gateway translates between the two formats. But a request keeps all its provider features only when your surface matches the format of the route that serves it. Then the gateway forwards the request as you sent it. The table gives the native surface of each model family's first route: | Model ids | Native surface | Features that need it | |---|---|---| | `claude-*` | `/v1/messages` (Anthropic SDK) | Prompt caching with `cache_control`, `thinking`, server tools, image input | | `gpt-*` | `/v1/chat/completions` (OpenAI SDK) | `response_format`, `reasoning_effort`, `seed`, `logprobs`, image input | | `author/model` ids, for example `deepseek/deepseek-v4-flash` | `/v1/chat/completions` (OpenAI SDK) | The same OpenAI fields, where the host supports them | On the other surface, the gateway translates. A field that the translation cannot carry is dropped and named in the `x-tm-dropped-params` response header. If dropping the field would change the answer, the request is a 400 that names the field instead. Before it refuses, the gateway tries another provider of the same model that speaks your format. The `x-tm-provider` header names the provider that served you. The catalog gives each model's native format in `providers[0].dialect`. If your code sends only text and function tools, either surface is fine for every model. The native surface helps only while the first route serves you. A Claude request on `/v1/messages` that is pinned to `openrouter`, or fails over to it, is translated: `thinking`, server tools and image blocks cannot cross (the request is a 400 if no other route can take it), and other Anthropic-only fields such as `top_k`, `metadata` and `cache_control` are dropped and named. Details: [Wire compatibility](compat.md). ## 3. Call models ### Rules for every request 1. **Model ids are exact.** Copy them from `GET /v1/models` or `/api/models.json`. Claude and GPT ids are bare: `claude-sonnet-5`, `gpt-4o-mini`. Open models keep their author prefix: `moonshotai/kimi-k3`. OpenRouter variants such as `:free` do not exist here. A dated snapshot of a listed model, such as `claude-haiku-4-5-20251001`, also works: it routes and bills as its model. 2. **Keep model ids in configuration**, in one place. Check them against `GET /v1/models` when your service starts. Then a model change is a configuration change. 3. **Use the right route for the output.** Chat models answer on the two chat surfaces. Image models answer only on `POST /v1/images/generations`, video models only on `POST /v1/videos`, and decision models only on `POST /v1/decisions`. The catalog field `output_modalities` (`text`, `image`, `video` or `decisions`) tells you which is which. A decision model also takes its own question format: the catalog field `decisions_wire` in `/api/models.json` names it (`systemone` for Jev, Mercury Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider v1 27B and Sage, and `gliner`), and [POST /v1/decisions](/docs/api-decisions) gives each one. Do not guess the format from the model's name. 4. **Send `max_tokens` every time**, from 1 to 32,768. If you leave it out, the gateway writes `max_tokens: 4096` into the request, and a longer answer stops with `finish_reason: "length"`. A value above 32,768 is a 400. The gateway also holds balance for `max_tokens` while the call runs, so a realistic value helps on a small balance. 5. **Ask for one answer.** `n` above 1 is a 400. Make separate calls for several samples. 6. **Send the whole conversation on every turn.** The gateway keeps no conversation state, and it stores no prompts or answers. 7. **Keep `temperature` at 1 or less** in code that can reach Claude models. On a translated request, a higher value is a 400. 8. **Refuse drops where they matter.** A field outside the forwarded set is dropped and named in `x-tm-dropped-params`. Send `"provider": {"require_parameters": true}` to get a 400 instead of a drop. The 400 comes from the first route that would drop a field: the gateway does not go on to a later route that could carry it. So combine it with `only` or `order`. For example, the `anthropic` route drops `seed`, so for `seed` on a Claude model over `/v1/chat/completions`, send `{"only": ["openrouter"], "require_parameters": true}`. ### Streaming Stream answers that a person watches, and long answers. The OpenAI surface sends, in this order: a role chunk, content chunks, a finish chunk, a usage chunk with an empty `choices` list and `usage.cost`, then `data: [DONE]`. The Anthropic surface sends the native event sequence, with usage and `cost` in the final `message_delta`. - If the provider fails after the answer starts, the stream ends with one error event and nothing after it (no `[DONE]`). The SDKs raise an error. The partial answer is billed. The gateway never continues an answer on another provider. To try again, send the whole turn again. - During a silence of 15 seconds or more, the gateway sends a keep-alive: an SSE comment on the OpenAI surface, a `ping` event on the Anthropic surface. The SDKs skip them. If you parse SSE yourself, skip them too. - The first response headers must arrive within 20 seconds, or the gateway tries the next provider. A request can run for at most 450 seconds. - To stop an answer, close the connection. The provider stops. You pay for the usage that the provider reported. If it reported none, you pay for an estimate of the text already streamed, at about four characters per token. If no text streamed yet, the gateway sees no usage, and you pay the full hold for the call: the input estimate plus `max_tokens`, at the model's prices. Python — the OpenAI SDK, keeping the cost and handling a failure in the middle of an answer: ```python from openai import APIError parts, usage = [], None try: stream = client.chat.completions.create( model="claude-sonnet-5", max_tokens=2048, stream=True, messages=messages, ) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: parts.append(chunk.choices[0].delta.content) if chunk.usage: # the usage chunk; when usage is known, it also comes before an error usage = chunk.usage except APIError: # The answer failed part way. What arrived is billed. Do not join a retry onto # `parts`: send the whole turn again, or show the error. raise print("".join(parts), "cost:", usage.cost if usage else None) ``` Details: [Streaming](streaming.md). ### Calls that do not stream Set the client timeout above 450 seconds. The SDK default of 10 minutes is fine. The Anthropic SDKs refuse a call that does not stream when `max_tokens` is above about 21,000, unless you pass an explicit `timeout`. Stream those calls instead. Do not close the connection of a call that does not stream. The gateway stops the provider, but it sees no usage, so you pay the full hold for the call: the input estimate plus `max_tokens`, at the model's prices. If you may need to cancel, stream the call. An image render is different: it does not stop. It finishes and is billed at the usage that the provider reports. ### Tools, reasoning and structured output - Function tools work on both surfaces for every chat model. The gateway translates tool calls and tool results. - Anthropic server tools (web search, computer use, bash, text editor) work only for Claude models on `/v1/messages`, while the `anthropic` route serves them. The `anthropic` route refuses `mcp_servers` and `container`. The request then falls over to `openrouter`, which runs it without them and names them in `x-tm-dropped-params`. Send `"provider": {"require_parameters": true}` to get a 400 instead. - `thinking` works only for Claude models on `/v1/messages`, while the `anthropic` route serves them. Reasoning comes back in different fields. On `/v1/chat/completions`, Claude's thinking arrives as `reasoning_content`, and open models return OpenRouter's own `reasoning` field (and `reasoning_details`). On `/v1/messages`, a streamed answer from a non-Claude model carries its reasoning as `thinking` blocks, and a non-streamed answer does not carry it. You can send `thinking` blocks back in the next turn. - `response_format` needs an OpenAI-format provider: GPT and `author/model` ids. For JSON from a Claude model, force a tool call with `tool_choice` and read the tool input. - Do not end the message list with an assistant message when a Claude model can serve the request from the OpenAI surface. That is a 400, because Claude rejects a prefilled answer. ### Images in prompts Send image parts on the model's native surface: Claude on `/v1/messages`, GPT and `author/model` ids on `/v1/chat/completions`. When the gateway must translate, an image part is a 400 that names the part. Send an image, document, audio or video part only to a model that takes it: `architecture.input_modalities` in `GET /v1/models` lists each model's inputs. A part the model does not take is a 400 that names the part, and nothing is billed. Each image or document part counts as 65,536 tokens against your limits and your balance hold while the call runs. ### Other routes - **Images:** `POST /v1/images/generations` takes the OpenAI Images API body. The image comes back as base64 in the response (`response_format: "url"` is a 400). `n` is 1 unless the model allows more: see `supported_parameters.n` in `/api/models.json` (GPT Image models allow up to 4). There is no streaming. A failed render bills nothing. See [POST /v1/images/generations](api-images.md). - **Video:** `POST /v1/videos` starts a job. Poll `GET /v1/videos/{id}` until it ends, then download `GET /v1/videos/{id}/content`. A failed job costs nothing. See [POST /v1/videos](api-videos.md). - **Decisions:** `POST /v1/decisions` sends a state and typed questions to a decision model: `typesafe/jev-1.13` (Jev, from TypeSafe), `inception/mercury-decide` (Mercury Decide, from Inception), `bespokelabs/nimble-v3` (Bespoke Nimble v3, from Bespoke Labs), `cloudflare/clef` and `cloudflare/clef-flash` (Clef and Clef-flash, from Cloudflare), `routerplus/decider-2b` and `routerplus/kev-4b` (Decider 2B and Kev 4B, from RouterPlus), `perplexity/pplx-decider-v1-27b` (Perplexity Decider v1 27B, from Perplexity), `levanto/sage-1.2` (Sage, from Levanto) or `fastino/gliner-2.5-decide` (GLiNER-2.5-Decide, from Fastino). It returns one typed answer for each question, such as a choice, a score or a yes/no probability. The question format follows the model: System One questions (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage) have a `type`, GLiNER takes a `schema` of classifications and extractions in place of `questions`, and a question in another model's format is a 400. GLiNER's numbers are confidences, not calibrated probabilities. Perplexity Decider bills the state once per question, so on it ask only the questions you need. The gateway never fails over from one model to another. Use it for routing, classification and moderation points in your code, in place of a chat model that you parse. See [POST /v1/decisions](api-decisions.md). - **Token counting:** `POST /v1/messages/count_tokens` is free, for models that have an Anthropic-format provider. It still counts against your requests per minute. - **Your deployed models:** an id such as `tm/-v1` from [Optimize](model-search.md) works on `/v1/chat/completions` only, with keys of the organization that deployed it. With `stream: true` it returns the finished answer as one chunk. ## 4. Pin providers (only when you must) How routing works without a pin: - Each chat model has a fixed, ordered list of providers. Today Claude models go to `anthropic` first and `openrouter` second. GPT models go to `openai` first and `openrouter` second. Open models (`author/model` ids) go to `openrouter` only. - The order does not change from request to request. There is no random spread across providers. - The gateway moves to the next provider only when the current one fails before it sends any output, or has no capacity left. - OpenRouter picks the host that runs an open model, for example Fireworks, Together or DeepInfra, for each request. - The price of a model is the same whichever provider serves it. Pin when a rule requires one provider, when a feature exists at one provider only, when a session must stay on one host (step 5), or when you compare providers in a test. A pin can cost you fallbacks: `only` and `"allow_fallbacks": false` remove the fallback provider, and `upstream` removes the fallback hosts inside OpenRouter. `order` keeps every fallback. ### The `provider` object Send it in the request body on `/v1/chat/completions` or `/v1/messages`. The gateway reads it and never forwards it. An unknown key inside it is a 400. | Field | Value | Effect | |---|---|---| | `order` | provider ids | Try these providers first, in this order. The others stay as fallbacks. | | `only` | provider ids | Use only these providers. | | `ignore` | provider ids | Never use these providers. | | `allow_fallbacks` | boolean | `false` keeps only the first route. If that route is resting after failures, the request fails with 503 and `retry-after` instead of moving on. | | `upstream` | host tags, up to 8 | On an OpenRouter route: use only these hosts, in this order, and never another host. If they all fail, the request fails. | | `require_parameters` | boolean | `true`: when the first route that would serve you would drop one of your request fields, the whole request is a 400. The gateway does not try a later route. | | `connections`, `connection_order`, `funding` | ids, `"byok"` or `"house"` | Choose among your own provider connections and who pays. See [Routing policies](routing-policies.md). | | `regions`, `zdr`, `data_collection` | region list, boolean, `"allow"` or `"deny"` | Hard filters. A route passes only with operator evidence, and marketplace routes have none today. So on marketplace traffic, `regions`, `"zdr": true` or `"data_collection": "deny"` removes every route. | `order`, `only` and `ignore` take route labels and OpenRouter host tags. The marketplace routes are `anthropic`, `openai` and `openrouter`: the `providers[].id` values in `/api/models.json`. Your own connections use their profile: `openai`, `anthropic`, `azure` or `bedrock`. - A label that matches a route label selects or orders that route, even when that route does not serve the model. Route labels win over host tags with the same spelling: `only: ["anthropic"]` means the Anthropic route, so on an open model it leaves no route. - Any other label is a host tag. When the `openrouter` route runs, the gateway sends the tags to OpenRouter as its own `order`, `only` and `ignore`, with your `allow_fallbacks` (default true). OpenRouter then prefers, limits or skips those hosts, and it can still fall back to other hosts unless `allow_fallbacks` is `false`. - A host tag in `only` keeps the `openrouter` route open and closes the routes that `only` does not name. - `upstream` is the strict pin: those hosts only, in that order. Host tags from `order`, `only` and `ignore` do not go with it. - `POST /v1/route` shows the host preferences in the field `openrouter_provider` (`null` when there are none). Host tags are OpenRouter's provider slugs, for example `fireworks`, `together`, `deepinfra`, `amazon-bedrock` or `google-vertex`. The hosts that run a model are its `served_by` entries in `/api/models.json`. Use only those hosts: the gateway tells OpenRouter to skip some hosts that OpenRouter lists, and a request that names only those hosts fails. For most hosts, `served_by[].id` is the same as the tag. Three differ: `amazon` (tag `amazon-bedrock`), `moonshot` (tag `moonshotai`) and `zai` (tag `z-ai`). | Goal | `provider` | |---|---| | Claude only from Anthropic's own API | `{"only": ["anthropic"]}` | | Claude through OpenRouter first, Anthropic as the fallback | `{"order": ["openrouter"]}` | | Claude on Amazon Bedrock | `{"only": ["openrouter"], "upstream": ["amazon-bedrock"]}` | | An open model on one host for every turn | `{"upstream": ["fireworks"]}` | | An open model on one host, with one named backup host | `{"upstream": ["fireworks", "together"]}` | | Fail instead of falling back to another provider | add `"allow_fallbacks": false` | Check a pin before you ship it. `POST /v1/route` returns the ordered candidates and each excluded route with its reason. It calls no provider and costs nothing. If your controls exclude every route, a real request returns 404 `model_unavailable`, the same error as an unknown model. So when a pinned request gets a 404, run `/v1/route`. ```bash curl -s https://api.routerplus.com/v1/route \ -H "Authorization: Bearer $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"claude-sonnet-5","provider":{"only":["openrouter"],"upstream":["amazon-bedrock"]}}' ``` Python — the OpenAI SDK sends fields it does not know through `extra_body`: ```python reply = client.chat.completions.create( model="moonshotai/kimi-k3", max_tokens=1024, messages=[{"role": "user", "content": "Hello"}], extra_body={"provider": {"upstream": ["fireworks"]}}, ) ``` TypeScript — pass a variable, so the field the SDK types do not know is kept: ```typescript const request = { model: "moonshotai/kimi-k3", max_tokens: 1024, messages: [{ role: "user" as const, content: "Hello" }], provider: { upstream: ["fireworks"] }, // read by the gateway, never forwarded }; const reply = await client.chat.completions.create(request); ``` Check each answer with its response headers: `x-tm-provider` names the provider, `x-tm-served-by` names the host on an OpenRouter route as OpenRouter spells it (for example `Amazon Bedrock`, not the tag `amazon-bedrock`), `x-tm-upstream-model` gives the id the provider received when it differs from yours, and `x-tm-attempts` counts provider attempts. `GET /v1/generation?id=` shows the route of every attempt. To pin for a whole organization, workspace or key instead of each request, save a routing policy with the policy API. It takes a console session, so a human sets it up. A request's `provider` object can narrow a saved policy, but it can never add routes. See [Routing policies](routing-policies.md). ### Your own provider keys (BYOK) You can route through your own OpenAI, Anthropic or Azure key, or an AWS Bedrock role. A human adds it once in the console as a connection. Your code still calls the gateway with a marketplace key. The provider bills you directly, and `usage.cost` is 0. A key bound to a connection can call only that connection's exact model ids, and it does not fall back to marketplace supply unless a routing policy allows it. See [Bring your own key](byok.md). ## 5. Keep sessions on one route and caches warm A prompt cache saves money and time only when the next request reaches the same provider, or the same host, with the same prompt start. This is how the gateway treats sessions today: - The gateway keeps no session state. Your app stores the conversation and sends all of it on every turn. - Routing is deterministic. For the same model, key and `provider` object, every request tries the same route first. So the turns of a session stay on one provider without a session id. - A turn moves to the next route when the first route fails before it sends output (one failure is enough), when the route's capacity pool is full, or while its provider's 429 pause lasts (the provider's `retry-after`, 1 to 60 seconds). After two failures in a row, later turns skip the route for 30 seconds. A turn that moves can miss the cache once. On a follow-up turn, the gateway first waits up to 5 seconds for a full pool: see [Waiting for the first route](#waiting-for-the-first-route). - Send a session id, and the gateway gives it to each provider in the provider's own field: see [Sessions and caching](#sessions-and-caching). Claude Code's `x-claude-code-session-id` header counts, with no change on your side. - For an open model, OpenRouter picks the host. With a session id (or the id that the gateway makes when you send none), OpenRouter keeps the conversation on one host from its first request, but it does not promise to. To force one host, pin it with `upstream`, as shown below. Remember that `upstream` turns off host fallback: if the pinned hosts fail, the request fails. Rules for a high cache hit rate: 1. Keep the start of the prompt byte-identical: the same system prompt, the same tool definitions in the same order, and earlier messages unchanged. Only append. 2. Put content that changes (the time, ids, documents retrieved for this turn) after the stable part, at the end. 3. Keep the model and the `provider` object the same for the whole session. 4. For Claude, put `cache_control` on the block that ends the stable part (a system, tool or message block), or send the top-level `cache_control` field for automatic caching. Both surfaces keep the marks for Anthropic and OpenRouter: see [Cache marks](#cache-marks). On marketplace routes the cache lasts 5 minutes: `ttl: "1h"` is removed. 5. OpenAI caches long repeated prompt starts by itself. Many open-model hosts do too. You set nothing for them. 6. For an open model, pin one host for each session when cache hits matter more than host fallback. 7. Measure. Cache reads are `usage.prompt_tokens_details.cached_tokens` on the OpenAI surface and `usage.cache_read_input_tokens` on the Anthropic surface. Cache writes are `cache_write_tokens` and `cache_creation_input_tokens`. Cache reads are billed at the lower `cached_prompt` price. Claude cache writes are billed at the higher `cache_write` price. Prices are in `/api/models.json`. Python — the Anthropic SDK, with a cache mark on a long system prompt that does not change: ```python reply = client.messages.create( model="claude-sonnet-5", max_tokens=1024, system=[{ "type": "text", "text": STABLE_INSTRUCTIONS, # the same bytes on every turn "cache_control": {"type": "ephemeral"}, # cache everything up to here }], messages=history, # append only ) print(reply.usage) # cache_read_input_tokens, cache_creation_input_tokens and cost ``` The provider caches a prompt start only above a minimum length, so a short system prompt shows no cache reads. For open models, give each session its own host. The same session always gets the same first host, and different sessions spread across hosts. The second host is used only when the first one fails: ```python import hashlib # Hosts that run the model: its served_by entries in https://app.routerplus.com/api/models.json (tags as in step 4) HOSTS = ["fireworks", "together", "deepinfra"] def session_route(session_id: str) -> dict: i = int(hashlib.sha256(session_id.encode()).hexdigest(), 16) % len(HOSTS) return {"upstream": [HOSTS[i], HOSTS[(i + 1) % len(HOSTS)]]} reply = client.chat.completions.create( model="moonshotai/kimi-k3", max_tokens=2048, messages=history, extra_body={"provider": session_route(session_id)}, ) print(reply.usage.prompt_tokens_details.cached_tokens) ``` ### Sessions and caching #### Session ids On `POST /v1/chat/completions` and `POST /v1/messages`, the session id is the first valid value among these inputs, in this order. Token counting, images and videos do not use it. | Order | Input | Kind | |---|---|---| | 1 | `session_id` | Top-level body field | | 2 | `x-session-id` | Header | | 3 | `x-session-affinity` | Header | | 4 | `x-claude-code-session-id` | Header that Claude Code sends by itself | - A valid value has 1 to 256 printable ASCII characters. `prompt_cache_key`, a top-level body field, follows the same rule. - The gateway ignores an invalid value and uses the next input. An invalid value alone is never an error. The ignored input is named in `x-tm-dropped-params` as `session_id (invalid)`, `prompt_cache_key (invalid)`, `header:x-session-id (invalid)`, `header:x-session-affinity (invalid)` or `header:x-claude-code-session-id (invalid)`. With `"require_parameters": true`, that makes the request a 400, as any dropped field does. - Valid session inputs are never named in `x-tm-dropped-params`, and they never trigger `require_parameters`. - `prompt_cache_key` stays OpenAI's own field. It wins over the session id on OpenAI routes, and it is the fallback on OpenRouter. - When two valid inputs disagree, the order decides. There is no error. - Codex's `session-id` header is not read. It can matter only after the gateway adds `POST /v1/responses`, because current Codex calls only the Responses API. What each route receives: | Route | Field | Value | |---|---|---| | `openai` (marketplace), and `openai` and `azure` (your connections) | `prompt_cache_key` | Your `prompt_cache_key`, or else the session id. Nothing if you sent neither | | `openrouter` (marketplace) | `session_id` | The session id, or else your `prompt_cache_key`, or else an id that the gateway makes. If you sent `prompt_cache_key`, it goes too | | `anthropic` and `bedrock` | Nothing | Anthropic has no session input. Its cache matches the prompt start | - Values go to the provider unchanged. A value longer than the provider accepts is replaced by `tm-` followed by 43 base64url characters: a hash of the value. - For OpenRouter only, when you send neither a session id nor `prompt_cache_key`, the gateway makes an id from your organization id, the model and the opening messages, up to the first user message. OpenRouter then keeps the conversation on one host from its first request. - The gateway keeps no session table. The id goes to the provider, and the provider routes on it. - `prompt_cache_retention` goes unchanged to `openai` and `azure` routes, both `"in_memory"` and `"24h"`. On other routes it is dropped and named. #### End-user ids on marketplace routes On marketplace routes, the provider gets a hash in place of your end user's id, and the raw values do not go. The id is the first string among the body fields `safety_identifier`, `user` and `metadata.user_id`. The hash is `tm-` followed by 43 base64url characters, made from your organization id and that id. Without an id, it is made from your organization id alone. Each marketplace route gets exactly one field: | Marketplace route | Field that carries the hash | Fields removed | |---|---|---| | `openai`, `azure` | `safety_identifier` | `user` | | `openrouter` | `user` | `safety_identifier` | | `anthropic` | `metadata.user_id` | `user`, `safety_identifier` | | every other route (OpenAI-compatible), dedicated capacity included | `user` | `safety_identifier`, `metadata.user_id` | - A provider that blocks an id then blocks one user of one customer, not the whole marketplace account. - On your own connections, `user` and `metadata` pass unchanged. `safety_identifier` passes unchanged to `openai` and `azure` connections, and it is dropped and named on `anthropic` and `bedrock` connections. - Identity fields that a route uses are never named in `x-tm-dropped-params`. #### Cache marks The gateway never adds a cache mark that you did not send. Block and part marks (`cache_control`) are handled this way: | Request | Route | What happens to the marks | |---|---|---| | `/v1/chat/completions` | Anthropic direct or Bedrock | Each text part keeps its mark. A marked system part turns the system prompt into blocks. A tool message's last mark goes on its `tool_result` block. Other part keys are not forwarded | | `/v1/chat/completions` | OpenRouter | Parts pass unchanged | | `/v1/chat/completions` | OpenAI direct or Azure | Part marks pass unchanged. OpenAI and Azure ignore them | | `/v1/messages` | Anthropic direct or Bedrock | Block marks pass unchanged | | `/v1/messages` | OpenRouter | Marks survive in system, user, assistant and tool content. Marks on tool definitions and on `tool_use` blocks are removed and named as `tools[i].cache_control` or `messages[i].content[j].cache_control` | | `/v1/messages` | OpenAI direct or Azure | Marks are removed and named | A top-level `cache_control` field (automatic caching), on either surface: - Anthropic direct and OpenRouter: sent unchanged. - Bedrock: turned into a mark on the last block that can carry one, counted back from the end of the messages. If no block can carry it, it is dropped and named as `cache_control`. - OpenAI direct and Azure: dropped and named as `cache_control`. - Any route: a field that Anthropic would refuse is dropped and named as `cache_control`. Anthropic refuses a field other than `{"type": "ephemeral"}` with an optional `ttl` of `"5m"` or `"1h"`, a fifth mark (four marks exist already), a 1-hour field after a 5-minute mark, and a field whose TTL differs from the mark on the last block. On marketplace routes, `ttl: "1h"` is removed from every mark, so the cache lasts 5 minutes. The drop is named as `.cache_control.ttl`, for example `system[0].cache_control.ttl`, `messages[2].content[0].cache_control.ttl`, or `cache_control.ttl` for the top-level field. `ttl: "5m"` passes. Your own connections keep `ttl: "1h"`. #### Waiting for the first route On a follow-up turn (an assistant or tool message after a user message), if the first route's capacity pool is full, the gateway waits for that pool for up to 5 seconds (an operator setting) before it moves the request to the next route. It never waits past the request's deadline. The `x-tm-affinity-wait-ms` response header gives the wait in milliseconds, only when the wait was more than 0. ## 6. Handle errors and retries Every failure carries one stable class, `error_type`. Read it from `error.metadata.error_type` on the OpenAI surface, or `error.error_type` on the Anthropic surface. The `x-tm-error-code` header carries it too, but browsers cannot read that header. Never parse the message text. Log the `x-request-id` header with every failure. | `error_type` | HTTP | Retry | What to do | |---|---|---|---| | `rate_limit` | 429 | Yes, once | Wait `retry-after` seconds. If `x-tm-limit-kind` is `concurrency`, send fewer requests at once. | | `email_not_verified` | 403 | No | Stop and tell a human to open the verify link in their email; the key starts working within seconds of the click. | | `insufficient_quota` | 429 | No | Stop and tell a human: add credits or raise a spend cap. On a cap, `x-tm-cap-reset` gives the reset time. | | `upstream_error` | 5xx | Yes, once, after about 2 seconds | The gateway already tried every other provider it could. If the retry fails too, report it with the `x-request-id`. | | `upstream_error` | 4xx | No | The provider refused the request itself. Read the message. | | `gateway_error` | 503 | Yes | Wait `retry-after`. The request did not reach a provider. | | `gateway_error` | 504 | Yes, once | The deadline passed while the gateway tried providers. An earlier attempt may be billed: check `GET /v1/generation?id=` first. | | `gateway_error` | 400 | No | Send one output limit, from 1 to 32,768. | | `gateway_error` | 500 | No | Report it with the `x-request-id`. | | `model_unavailable` | 404 | No | Fix the model id, or fix your `provider` object (check it with `POST /v1/route`). | | `model_unavailable` | 502 | Yes, once | The connection to every provider failed. | | `invalid_request` | 400 | No | Fix the field that the message names. | | `context_overflow` | 400 | No | Shorten the input, or use a model with a larger context. | | `content_policy` | 400 | No | Change the prompt. It is never sent to another provider. | | `auth` | 401 | No | Check the header, the key, and whether the key was disabled. A 401, 402 or 403 relayed from every provider is the marketplace's own account problem: report it with the `x-request-id`. | | `request_too_large` | 413 | No | Keep the body under 10 MB (1 MB for a video request). | | `not_found` | 404 | No | Use a route that exists. The list is in [Errors](errors.md). | Rules: 1. Let the SDK retry. The official OpenAI and Anthropic SDKs retry 429 and 5xx responses twice by default, and they wait for `retry-after`. Keep that default. Do not add a second retry loop around it. The SDKs also retry an `insufficient_quota` response. That costs nothing, but your code must then stop. 2. There is no idempotency key. A retried request is a new request. If it reaches a provider, it is billed again. A refusal by the gateway's own limits (`x-tm-error-origin: gateway_admission`) cost nothing. 3. Never retry `insufficient_quota`, `invalid_request`, `context_overflow`, `content_policy` or `auth` in a loop. The request, or the account, must change first. 4. An error in the middle of a stream is final for that answer. Send the whole turn again. Never join two partial answers. Python — the OpenAI SDK, reading the class: ```python from openai import APIStatusError try: reply = client.chat.completions.create(model=MODEL, max_tokens=1024, messages=messages) except APIStatusError as e: error = e.response.json().get("error") or {} kind = (error.get("metadata") or {}).get("error_type") or e.response.headers.get("x-tm-error-code") request_id = e.response.headers.get("x-request-id") if kind == "insufficient_quota": notify_operator(request_id, error.get("message")) # money or a cap: a retry cannot help raise ``` Details: [Errors](errors.md). ## 7. Limits, spend caps and your users Default limits: | Scope | Requests per minute | Tokens per minute | In flight at once | |---|---|---|---| | Trial key | 20 | 1,000,000 | 8 | | Console or identity API key | 300 | 1,000,000 | 8 | | Organization, all keys together | 300 | 1,000,000 | 8 | Rules: 1. Limit how many requests your code sends at once. By default the organization allows 8 in flight. A ninth gets 429 `rate_limit` with `x-tm-limit-kind: concurrency`. Use a semaphore or a fixed pool of workers. To run more at once, a human first raises the organization and key limits in [/console/limits](https://app.routerplus.com/console/limits). A key's requests per minute cannot go above its ceiling (300 for a console key). Each marketplace provider pool also limits one organization to half of the pool's capacity. That refusal carries `x-tm-limit-scope: pool`. 2. Read `x-tm-remaining-rpm` and `x-tm-remaining-tpm` on each response, and slow down before they reach zero. `GET https://api.routerplus.com/v1/limits?model=` returns the limits that apply to your key. 3. While a call runs, token limits count an estimate: one token for each four bytes of the request, rounded up, 65,536 for each image or document part, plus `max_tokens`. An image sent inline as base64 counts its 65,536 only, not its bytes. The real count replaces the estimate when the call ends. Small requests and a realistic `max_tokens` let more calls run at once. 4. Spend caps are monthly, for each key and for the organization, on the UTC calendar month. A human sets them in the console. A cap refuses with 429 `insufficient_quota` and the `x-tm-cap-reset` header. Give each feature or customer its own key and cap, so a bug or a loop cannot spend everything. Your users: - By default, one backend key serves all your users. The gateway does not know who your users are, so enforce per-user quotas in your app. - `safety_identifier`, `user` and `metadata.user_id` identify your end user to the provider. On marketplace routes the gateway sends only a hash of that id with your organization; on your own connections they go unchanged (see [End-user ids on marketplace routes](#end-user-ids-on-marketplace-routes)). They do not identify, limit or bill anyone on the gateway. Do not put emails or names in them. - For limits that the gateway enforces per user, use the identity API. Create an `end_user` principal for each user, issue a key for it, and set limits on the principal. Your backend keeps those keys. Never give a key to the user. The identity API takes a console session, not an API key, so a human with an owner or admin role sets it up. See [Workspaces and principals](identity.md). - To name your app in analytics, send the `HTTP-Referer` and `X-Title` headers. They carry no content. Details: [Rate limits and spend caps](limits.md) and [Limits and capacity](admission.md). ## 8. Track cost - Every billed response has `usage.cost` in USD. It is the full charge for the request, including provider attempts that failed before the answer. Store it with the `x-request-id`, the model and your own ids, such as the user and the feature. - In a stream, `cost` is in the usage chunk on the OpenAI surface, and in the last `message_delta` on the Anthropic surface. - An Anthropic-surface stream that fails in the middle of the answer has no usage in it. Read its cost from `GET /v1/generation?id=`. - `GET /v1/usage` returns the balance and the last 20 attempts. For a key bound to its own workspace or principal, the balance fields are `null`. The console shows usage and logs for the whole organization. - BYOK calls show a cost of 0, because your provider bills you directly. - To check a charge, multiply the token counts by the prices in `/api/models.json`. The gateway rounds the final charge down. Details: [Pricing and billing](pricing.md) and [GET /v1/usage and /v1/generation](api-usage.md). ## 9. Verify the integration One cheap call shows the answer, the cost and the headers that the gateway adds: ```bash curl -sS -i https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"gpt-4o-mini","max_tokens":16,"messages":[{"role":"user","content":"Reply with OK"}]}' ``` Then confirm that the ledger saw it: ```bash curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usage ``` The integration is done when all of these are true: 1. Every client uses the gateway base URL and reads the key from `TM_API_KEY` on the server. 2. The key is not in the repository, in a client bundle or in a log. 3. Every model id in the configuration appears in `GET /v1/models`. 4. Every request sends `max_tokens`. 5. One streamed and one non-streamed test call return `usage.cost`. 6. The `x-tm-dropped-params` header is absent, or it names only fields that you accept losing. 7. `x-tm-provider`, and `x-tm-served-by` when you pin a host, show the route that you expect. (`x-tm-served-by` uses OpenRouter's display name, not the tag.) 8. The error handling switches on `error_type`. It stops on `insufficient_quota`, and it adds no retry loop on top of the SDK's own retries. 9. The number of requests in flight stays within the organization's limit. 10. The logs keep `x-request-id` and `usage.cost` for every call. 11. The test calls appear in `GET /v1/usage`. ## Leave instructions for the next agent Add this block to the `AGENTS.md` or `CLAUDE.md` file of the repository that you integrated. Later coding sessions then keep the integration correct: ```markdown ## LLM calls: RouterPlus gateway - Every LLM call goes through the RouterPlus gateway. Full guide: https://app.routerplus.com/docs/agent-integration.md - OpenAI SDK base URL: https://api.routerplus.com/v1. Anthropic SDK base URL: https://api.routerplus.com. Key: the TM_API_KEY environment variable, on the server only. - Model ids live in configuration. Each one must exist in https://app.routerplus.com/api/models.json. Never invent or alias an id. - Send max_tokens on every request (1 to 32768). Never send n above 1. - Claude models: /v1/messages, with cache_control on the stable start of the prompt. GPT and author/model ids: /v1/chat/completions. - Keep the start of the prompt byte-identical across turns. Only append. - Send one id per conversation as the x-session-id header (or session_id in the body). Providers then keep its cache warm. - Provider pins go in the request body "provider" object (only, ignore, order, allow_fallbacks, upstream, require_parameters). Check a pin with POST https://api.routerplus.com/v1/route. - Errors: switch on error_type. rate_limit: wait retry-after, then retry once. insufficient_quota: stop and tell a human. Never retry invalid_request, context_overflow or content_policy. - No idempotency key: a retry is a new, billed request. Add no retry loop on top of the SDK's own retries. - At most 8 requests in flight per organization, unless the limit was raised. - Log x-request-id and usage.cost for every call. - Not available: the Responses API, embeddings, audio, files and batches. ``` ## Steps only a human can do - Complete browser signup through Clerk, including email verification, then buy credits in Billing. Signup and card verification alone do not add credit. - Fix billing when `insufficient_quota` appears: add credits, or raise a spend cap in the console. - Sign in to the console to create keys, set spend caps and change limits. - Add your own provider keys (BYOK), save routing policies, and set up workspaces and end-user principals. - Approve large test runs. Every call is billed. ## Where next - [Quickstart](/docs/quickstart) — the first streamed call, step by step - [Errors](/docs/errors) — every error class, and what to do about it - [Routing policies](/docs/routing-policies) — the provider object, saved policies and BYOK routes - [Streaming](/docs/streaming) — frame order, keep-alives and failures in the middle of an answer - [Rate limits & spend caps](/docs/limits) — the numbers behind every 429 - [Wire compatibility](/docs/compat) — which fields pass, which translate and which drop --- Source: https://app.routerplus.com/docs/playground.md # Playground [`https://app.routerplus.com/playground`](https://app.routerplus.com/playground) is an interface over the same catalog, routing, prices and ledger as the API: chat, typed decisions, images and video. Sign in, pick a model, and each answer arrives with the provider that served it, the token counts, and what it cost. ## What it does - **Any catalog model.** The picker searches the catalog; `⌘K` / `Ctrl+K` opens it. It lists the models of the mode you are in first (chat models in Chat, decision models in Decision). Chat starts with GPT-6 Astra, Claude Opus 5.5, GLM 5.3 Flash, DeepSeek V4.1 Flash and Hy4 Preview; then the other open-weight models; then GPT-6.1 Sol, GPT-6 Luna and Claude Sonnet 5.5; then the rest, each group newest first. Decision mode lists Jev 1.13, Bespoke Nimble v3, Clef, Clef-flash, Mercury Decide, Kev 4B, GLiNER-2.5-Decide, Decider 2B and Sage 1.2, in that order. The search matches a model's name, id, lab, the providers that run it, and its kind: type `decision` for every decision model, `image` or `video` for those, `free` for the free ones. Each row carries the lab's mark, its price — per million tokens, or for an image or video model the provider's own price and the most one image or second can cost — and a chat model's context length. Typing an id that is not listed offers it verbatim, which is how a dated snapshot of a catalog model is pinned. - **Compare up to three chat models side by side** (Decision mode compares up to four). **Compare** adds a column. One prompt goes to every column at once, each streams into its own card, and each shows its own provider, tokens, latency and cost. Every column keeps its own thread: a follow-up sees that model's replies only, never another's. Removing a column leaves the rest untouched. - **Sample prompts** under the composer send a real prompt in one click. The first is always the **strawberry test** (how many times "r" appears in "strawberry"), asked plainly. Each reply to it says **passed**, **failed** with the count the model gave, or that it gave no clear count. The others are drawn at random: a reasoning trap, a code task, structured JSON, a table, a six-word story and a few more. They make way as soon as the conversation starts. - **Streaming replies** with a stop button, regenerate, a collapsible "thinking" section for reasoning models, and safe rendering of code blocks, lists, links and tables. A table is drawn only when the reply holds a real markdown table: a header row, then a divider row such as `|---|---|` with the same number of cells. Any other text with a pipe in it reads as it is. A wide table scrolls sideways inside its card. - **Code and Parameters** sit in the composer, to the left of Send. **Code** shows the request the playground is about to make, as curl, Python or Node, so you can paste it into your own project. **Parameters** holds a system prompt, temperature, top_p and max output tokens, per conversation. - **Cost on every reply:** the same `usage.cost` the API returns, plus time to first byte and an `audit` link to the request's ledger rows. - **Organizations:** if you belong to several, the sidebar selector switches which one the playground bills. ## Tetris Royale [Tetris Royale](https://app.routerplus.com/playground/tetris-royale) compares the eight current System One decision models on the same block sequence. Open it from the Playground mode bar; the arrow button at the top left goes back. It keeps the selected organization and uses a separate server-held managed key named **Tetris Royale**; ordinary Playground keys retain their existing limit. The server validates the board and match counters, then constructs the decision prompt and legal choices itself. The game relay does not accept custom prompts or questions. Every model request follows the organization's normal credits, limits, billing and provider routing. No API key reaches the browser. A match defaults to 200 pieces; Settings offers 20 up to a thousand. Each model plays independently: a slow or rate-limited model does not hold up the others. Transient 429s cool down only that model and retry up to three times, respecting Retry-After. Other errors offer an explicit retry. Concurrent matches in the same organization share server-side per-model pacing; a busy Mercury queue does not pause other models. Gateway quotas remain authoritative. Pause finishes current moves and stops new calls; switching tabs pauses the match. Reset starts over. Match state lives only in the open page. Settings offers five **Block sequences**: each is a repeatable order of pieces shared by all models. The boards keep a fixed order: Jev, Perplexity, Bespoke, Clef, Clef-flash, Mercury, Kev, Decider. **Standings** rank the models by score while the match runs. A row moves when its model passes another, equal scores share a place, and a line clear or a change of place shows for a moment. Standings compare scores at each model's current progress; final standings appear when every board finishes or tops out. Both modes use a 10 × 20 board and seven-bag pieces, line-clear and level points, hard-drop points, combos, back-to-back clears and all-clear bonuses. Settings → **Moves** picks how a piece reaches its place: - **Drop and slide** (the default): each piece turns to its orientation at the top, then falls; while it falls it may move left or right, so it can tuck sideways into a gap under an overhang. It never turns while falling. The model picks any final position the piece can reach this way, a straight drop or a tuck. Each choice tells the model which of the two it is and the points it scores now; the board shows the piece turn, drift toward its column and tuck. - **Drop only**: the model picks a rotation and a column, and the piece drops straight down. Hold, wall kicks, T-spins and a gravity deadline are not part of either mode; **Rules & scoring** states the exact rules. ## Decision Some models do not write text at all. A *decision model* takes a **state** — a message, a ticket, a list of items — and typed **questions**, and returns a typed answer for each with the probability or confidence behind it. It is built for routing, classification and the decision points inside an application, where a fast predictable answer matters more than prose. The catalog has ten: | | Jev 1.13 | Mercury Decide | Bespoke Nimble v3 | Clef and Clef-flash | Decider 2B and Kev 4B | Perplexity Decider v1 27B | Sage 1.2 | GLiNER-2.5-Decide | |---|---|---|---|---|---|---|---|---| | Catalog id | `typesafe/jev-1.13` | `inception/mercury-decide` | `bespokelabs/nimble-v3` | `cloudflare/clef`, `cloudflare/clef-flash` | `routerplus/decider-2b`, `routerplus/kev-4b` | `perplexity/pplx-decider-v1-27b` | `levanto/sage-1.2` | `fastino/gliner-2.5-decide` | | Made by | TypeSafe | Inception | Bespoke Labs | Cloudflare | RouterPlus, our own models | Perplexity | Levanto | Fastino | | Question format | System One | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | GLiNER's schema | | Question kinds | choice, rating, yes/no; tags as one yes/no per tag | the same as Jev | the same as Jev | the same as Jev | the same as Jev | the same as Jev | the same as Jev | yes/no, choice, rating, tags, all as labels | | Rating scale | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | exactly 5 levels | | Numbers it returns | probabilities | probabilities | probabilities, at full precision | probabilities, to 4 decimals | probabilities | probabilities, at full precision | probabilities | confidences, not calibrated probabilities | | When unsure | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best label, with a confidence | | Price | $0.042 per million input tokens | free while OpenRouter serves only its free variant | $0.04 per million input tokens | $0.24 (Clef) or $0.09 (Clef-flash) per million input tokens | free: our own models, at $0 | $0.04 per million input tokens, the state billed once per question | $0.05 per million input tokens, the state billed once per question | $0.03 per million input tokens | Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage are the nine **System One** models: they take the same questions and answer in the same shape, so wherever this page says what System One takes, it means all nine. Nimble, Clef and Clef-flash have two limits of their own: a choice needs at least 2 options, and one request holds at most 64 questions, each tag counted as one. Clef and Clef-flash have a third: every answer key is 1 to 100 letters, digits, `_`, `.` or `-`. A question id in Decision mode always fits that rule. A tag can break it: Decision mode asks each tag under the key `__`, and a long question id with a long tag id can pass 100 characters. Decision mode checks these limits before a run. It counts the questions in order, and only the ones the column will send: a question that would cross 64 shows **Not askable on Bespoke Nimble v3** (or on Clef, or on Clef-flash) and is left out of that column's request. It takes none of the 64, and a later question that still fits is sent. So the provider never refuses the column whole. For Decider 2B and Kev 4B, Decision mode checks no limit of their own: probed on 2026-10-02, each answers a one-option choice and a request of 200 questions, so Decision mode holds their columns to the family's rules alone, as it holds a Jev column. Perplexity Decider has one limit of its own: one request holds at most 128 questions, each tag counted as one. It takes a one-option choice, as Jev does. Decision mode counts the questions in order, as for Nimble, and a question that would cross 128 shows **Not askable on Perplexity Decider v1 27B**. Perplexity bills the state once for each question it asks, so on a Perplexity Decider column every question, and every tag of a tags question, costs one state. Switch to **Decision** and the composer becomes a decision console: a box for what the models judge, and up to 20 questions that every column is asked. Each question is a line of text and an answer type: **Yes / no**, **Pick one** (one of your options), **Rating** (a scale, lowest level first; a new rating starts with five levels) or **Tags** (which of your tags apply). Options, tags and levels are rows: a short name, which is what the model answers with, and what it means. You never write an answer key: Decision mode names each question from its words, and **Code** shows the names it sends. A yes / no saved by the earlier editor with descriptions of a yes and a no keeps them, shown under the question; only the System One models take them, and **Remove descriptions** makes the question plain again. An example loads by itself. The **Example** menu picks another, or **Start blank**. The examples are the jobs a decision model is for: checking an answer against its sources, reviewing an agent's run, catching a prompt injection, triage, bug severity, lead qualification, code review, moderation, feedback tags and a JSON state. **Run** runs it once. **Run 10 times** runs every column ten times (see below). ### One form for every model You write each question once. Each decision model has its own question format, so Decision mode translates the form into that format for each column. This is what each model receives: | Question | System One (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sage) | GLiNER-2.5-Decide | |---|---|---| | Yes / no | a `noul` question | a classification: the question text is the task, the labels are `yes` and `no` | | Choice, 2 to 20 options, descriptions optional | a `choice` question: each option, with its description or none | a classification: one label per option, `"name: description"` (or `"name"`), read back as the option's name | | Rating, exactly 5 levels, low to high | a `score` question with 5 levels | a classification: the labels `"0: "` to `"4: "`, read back as the level | | Tags, 1 to 20, descriptions optional | one `noul` question per tag, asked as ` Does this apply: ?` (with the tag's description after the tag, when it has one) | a multi-label classification that returns every tag with its confidence; a tag applies at 0.5 or more | GLiNER reads each label as text, so Decision mode puts an option's or level's description into its label. The description changes GLiNER's answer, as it would for a person. Decision mode asks GLiNER at threshold 0 (`cls_threshold: 0`), so GLiNER always names its top label, with its confidence beside it. At GLiNER's default of 0.5 a task answers `null` when no label reaches it; Decision mode shows a `null` as Not sure. System One has no tags question, so Decision mode asks a System One model one yes/no per tag. Each of those is a question in that column's request, and the column shows the tags together. This holds for every System One model alike (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sage): Decision mode builds a System One body from the catalog's word on the model's format, never from its name. ### Compare Put up to four models side by side with **Compare** — up to four of the ten decision models, or three of them and a chat model, or any mix; a chat, image or video conversation stops at three. One turn cannot hold all ten decision models: pick the four, or fewer, that you want to compare. A column's chip swaps one model for another. With two or more decision models in the comparison, the question builder holds to the table above: - A question in a form that one of the chosen decision models cannot take is marked before the run, with the models that can take it. Examples: a rating that does not have 5 levels (System One models only), a choice with one option (Jev, Mercury Decide, Decider 2B, Kev 4B and Perplexity Decider only: Nimble, Clef and Clef-flash refuse a one-option choice), more than 20 options, a question past the 64 keys of one Nimble, Clef or Clef-flash request or the 128 keys of one Perplexity Decider request, a tag whose key is longer than Clef's 100 characters. Models refused for the same reason share one line. The notices name the models in the comparison that take the question, so a Mercury Decide, Nimble, Clef, Decider 2B or Perplexity Decider column is named as itself, never Jev. - The run still sends each column the questions its model can take. A question that a column's model cannot take shows **Not askable on** that model in the column, and it is left out of that column's request. The rest of the questions still go. - GLiNER needs a different text for each question, because the question text is its task name, and a task name is at most 256 characters. Two questions with the same text, or a question longer than 256 characters, are not askable on GLiNER, and the notice says so. The examples use forms that every decision model takes, the tags example included. ### Run 10 times **Run 10 times** runs the whole comparison ten times — every column, the same state and questions — and shows one ledger instead of a card per column. Each column sends up to two requests at a time, and the ledger fills in as the runs come back: - For each model and question, the answer the model gave most often, and in how many of the runs it gave it (**10/10** when it never changed its mind). A pick-one question lists every option the runs chose, with how many runs chose it. - The averages of the model's numbers: the mean probability of yes, the mean confidence of a choice, the mean level of a rating (a marker on the track, with a faint dot for each run), and for tags how many runs each tag applied in. The bars and markers move to the new averages as each run lands. - At the head of each column, the mean time per run and the cost of all its runs, with each run's time as a dot on a strip and the mean as a line. - Each row says whether the decision models agree on their most common answers (for a rating, mean levels within 0.5 of each other). Each run is billed like a run: ten runs of four columns are 40 requests. A column that hits a rate limit stops and says so at its head; the other columns go on. **Stop** stops them all. On Decider 2B and Kev 4B, the ten runs send the same request, so RouterPlus answers the repeats from its cache, which keeps an answer for 600 s. Those runs come back sooner than the first one, cost $0, and repeat the first run's answer exactly: on these models the ten runs do not show how much the model's answer varies. The first request after a quiet period can take 15 to 20 s while RouterPlus starts the model. Below the columns, a compare view draws each question as one row, with one cell per column, in the same way for every model: | Question | Each cell shows | |---|---| | Yes / no | The verdict (Yes, No or Not sure) and a bar from 0 to 1. For the System One models the bar is the probability of yes. For GLiNER it is its confidence in the label it chose, labelled **confidence**. For a chat model, its stated answer. | | Choice | The chosen option (or Not sure) and its probability or confidence. Per-option bars where the model returns them: a System One model's add up to 1, GLiNER returns the chosen label only. | | Rating | One 0 to 4 track, with one marker per model: a System One model's score and GLiNER's chosen level, with the confidence beside it. | | Tags | A grid of tag by model: applies ✓, does not apply ✗ or not sure ?, with the score. | Each row says whether the decision models agree: the same verdict, the same option, a level within 0.5, or the same set of tags, with every tag answered (a tag one model is not sure about is nothing to compare, and the row says so). The view shows only numbers that a model returned, each labelled with what it means. It invents none. **GLiNER's numbers are confidences, not calibrated probabilities**: a GLiNER 0.9 and a Jev 0.9 do not mean the same thing. A Clef confidence and a Jev confidence do not mean the same thing either: Cloudflare computes Clef's in its own way, so compare their probabilities, not their confidences. Each column still shows its own cost and latency. A Mercury Decide, Decider 2B or Kev 4B column's cost is $0.00 with no list price or discount beside it: the model is free, so there is nothing to take off. ### Answer cards The answer cards show each model's own answer. A System One model's (Jev's, Mercury Decide's, Nimble's, Clef's, Clef-flash's, Decider 2B's, Kev 4B's, Perplexity Decider's and Sage's) show bars: the chosen option, the probability of each alternative, and the model's confidence. GLiNER's cards show the label it chose and its confidence, and for tags each tag's confidence. Add a chat model as a column and the same questions go to it in words, tags included where the prompt can say them. When it answers in the JSON it was asked for, its answers render as the same cards and in the compare view — so you can read the typed answer, the numbers, the latency and the cost side by side. **A Decision run is a real API call.** It goes through the gateway's [`POST /v1/decisions`](/docs/api-decisions) under the playground key, like a chat reply. It is billed to your organization at the model's price, counts against the playground key's limits and the organization's, and appears in `GET /v1/generation`. Each answer card has an `audit` link to the request's ledger rows. Decision models can carry a discount, which may make a run free: the cost line shows what you paid and, when a discount applies, the list price with the percent off. To call a model from your code, use [`POST /v1/decisions`](/docs/api-decisions) in that model's own question format. The translation above happens in Decision mode only; the API does not translate. ## Images Pick an image model — GPT Image, Nano Banana, Seedream and Grok Imagine are in the picker, marked **image** with the most one image can cost — and the composer describes an image instead of starting a chat. Under the description are the options that model declares: **size** or **shape**, **resolution**, **quality**, how many **images** (up to four, where the model allows more than one), **format** and **background**. A model that lacks an option does not show it. The estimate beside the send button is the most the render can cost: the per-image ceiling the gateway reserves. The bill is what the render actually cost, which is usually less. **Compare** puts up to three image models side by side. The same description goes to each, with the options it accepts. A comparison holds image models only: choosing a chat model to lead starts a new conversation, so text and pictures never share a history. A render usually takes 10 to 60 seconds, and each tile shows the time elapsed. A render cannot be stopped once it starts — the provider bills it either way — so there is no Stop button. It keeps going while you open another chat or start a new one: its chat shows a dot in the list until the picture arrives, and the picture lands there. **A render also survives a reload**, or a background tab the browser put to sleep: it runs on our side, and the page takes it up again when it comes back. A finished render waits there for 15 minutes; after that it is gone, though it was still billed. Each image has a **Download** button. The sample descriptions include the modality's known hard cases: exact spelling, counting, hands and a transparent background. A sample fills the description rather than rendering, because a render costs more than a chat reply. **Images stay in your browser.** They are kept in its own storage (IndexedDB), filed by your sign-in, the conversation and the reply, so a reload or a visit tomorrow shows them again. They are never uploaded. A render is metered exactly like [`POST /v1/images/generations`](/docs/api-images), under the playground key. ## Video Pick a video model — Seedance 2.5, MiniMax H3 Max or Wan 3.0, marked **video** with the most one second can cost — and the composer describes a video. Its options come from the model: **length**, **resolution**, **shape** and, where the model makes sound, **sound**. The estimate beside the send button is the most the turn can cost: each column's length at its model's per-second ceiling. The bill is what the provider charged, which is usually less. **Compare** puts up to three video models side by side, like images. The description goes to each with the options it takes: a length longer than a model makes becomes its longest, and a resolution it lacks becomes its own default. Each column's card says what it was asked for. A video takes from ten seconds to several minutes. Its tile says where the job is — **Queued at the provider**, then **Being made** — and for how long. You can leave: open another chat, reload, or close the tab and come back later. The job keeps going at the provider, and the page follows it again when you return. There is no Stop: the provider bills a job once it starts. When it is done the video plays in the card, with its cost and a **Download** button, and it is kept in your browser like an image. A video is metered exactly like [`POST /v1/videos`](/docs/api-videos), under the playground key. ## How it is billed Playground usage is metered exactly like an API call. The console mints one key named **Playground** per organization, tagged **managed** on [`/console/keys`](https://app.routerplus.com/console/keys) (turn on **Show playground keys** there to list it; the page hides managed keys by default); every reply is an attempt under that key, charged to the organization's credits, subject to its spend caps, and visible in the console and `GET /v1/generation`. BYOK connections bound to that key apply too. The key's value is never shown to anyone: the console holds it in memory only and calls the gateway on your behalf. The key is replaced every 24 hours and after a console restart; the one it replaces is retired. Disabling it from the console just makes the playground mint a new one on the next message. ## What is stored Nothing content-bearing on our side. Conversations live in your browser's local storage, up to 100 of them: use **Export** and **Import** to move them, **Clear all** to delete them. Images and videos live in your browser's IndexedDB storage, up to 1 GB before the oldest go first; deleting a chat, or **Clear all**, deletes its files too. Export carries the conversations, not the files — download the ones you want to keep elsewhere. The marketplace records only what it records for every request: model, provider, timestamps, token counts and cost. For a video it also keeps the job — its id, model, length, shape, status and charge — so the page can follow it; never the prompt or the file. See [Data policy](/docs/data-policy). ## Limits - No attachments. A chat takes text; an image or video model takes a text description. Image editing, image-to-video and reference images are not supported yet. - Up to three image renders in flight per organization at once — one per comparison column — and up to four images per render. A video is a job at the gateway: your balance must cover each column's reservation. - A conversation is capped at 400 000 characters per request and 200 messages, and a system prompt at 32 000 characters; start a new chat when a model reports a context overflow. - The managed key has its own request rate limit (60 per minute) on top of the organization's limits. A comparison spends one request per column, and the gateway holds a credit estimate for each: with a low balance, one column can be refused while another answers. A Decision run counts the same way, and **Run 10 times** sends ten requests per column, two at a time per column. - Columns served by the same provider share that provider's pool. There is room for a full comparison — three chat models, or the four columns in Decision mode — and for other people working at the same time; a comparison wide enough to exhaust it returns `rate_limit` on the later columns, whose card offers **Try again**. Mercury Decide's pool is small (20 requests a minute in all, 15 per organization, and 5 in flight at once, 3 per organization, because OpenRouter serves only its free variant), and OpenRouter caps free requests per day across our account: **Run 10 times** with Mercury Decide can meet either limit, and its column says **rate limited** when it does. Nimble's pool takes 7 requests at once in all, 5 per organization: Bespoke allows our whole account 8, and one is kept for our test environment. Clef and Clef-flash share one pool: 200 requests a minute and at most 8 at a time in all, 150 and 6 per organization. Decider 2B and Kev 4B share one RouterPlus pool, which is large (11,400 requests a minute and 59 in flight at once in all, 44 per organization), so your organization's own limits bind first. Perplexity Decider's pool takes 540 requests a minute and 8 in flight at once in all, 405 and 6 per organization: Perplexity allows our organization 10 requests a second. - Each model carries the mark of the lab that made it, in that lab's own colours, or the model's own mark where it has one. A model id the catalog does not list gets a neutral chip rather than a guessed logo. ## Errors Errors show the same class and remediation as the API's [error table](/docs/errors): a `rate_limit` asks you to wait, `insufficient_quota` points you at credits, `content_policy` is never rerouted to another provider. Every failed reply carries its request id. --- Source: https://app.routerplus.com/docs/model-search.md # Optimize [Optimize](https://app.routerplus.com/model-search) finds a cheaper model or mix that scores as well as yours on your own Braintrust evals, and deploys it as an endpoint: one model id you call like any other. ## Before you start 1. Connect Braintrust on [Integrations](https://app.routerplus.com/console/integrations). Paste a Braintrust API key. The key is checked with Braintrust and stored encrypted, and the page shows only its last four characters. Owners and admins can connect, replace or disconnect it; viewers can only look. (Langfuse is on the same page as **Request access**: it is set up with the team, not self-serve.) 2. In Braintrust, have one of these to search on: | Source | What the search uses | |---|---| | An **experiment** that ran on a dataset | The dataset is the test cases. The experiment's `metadata.model` is the baseline. Its prompt comes from `metadata.prompt_id` or `metadata.prompt_slug`, or the project's only prompt. | | A **dataset** | The rows are the test cases. You pick the prompt, and a baseline model if you want one. Without one, the strongest candidate becomes the baseline. | | **Production logs** | Each logged call is a test case, with the production reply as the expected turn. A candidate is scored on whether its reply is as good as production's. The baseline is the model most of the logged calls used, unless you pick one. | The baseline must be a catalog model or one of your own deployed endpoints (`tm/...`). A search takes up to 2,000 cases. For an experiment or a dataset, these parts of the rows matter: | Braintrust | Used as | |---|---| | A row's `input.messages` | One test of the next assistant turn. | | The project's scorers | Quality. A case's score is the mean of its scorers. | | Row `tags` | The slices in a result's detail view. | | Row `metadata.conversation_id` and `metadata.goal` | Keeps the turns of one conversation together. The goal is used for simulated customers. | Code scorers run in Braintrust. Prompt scorers (LLM graders) run through the gateway and are billed to your credits. ## Start a search Click **New search** at the bottom of the searches list. Pick the source, name the search, and set the **search budget**. The budget is the most the search spends on model calls, from $5 to $100. Model calls are billed to your credits as the search runs, under a managed key made for the search (turn on **Show playground keys** on `/console/keys` to see it, with a **managed** tag). Viewers cannot start or change a search. The search runs in stages. The budget pays for them in this order: 1. **Shortlist.** Public benchmarks pick the open models to try, as the Pareto frontier of benchmark score against price. Each model is tried on up to three of its hosts: the cheapest, the one at the highest precision, and the model maker's own. 2. **Pilot.** Every candidate answers a sample of the conversations. 3. **Dev.** The leaders answer all the dev conversations. 4. **Mixes.** Cascades, routers and votes are built from the saved answers; a search agent proposes mixes to try. These cost nothing extra. 5. **Held-out.** The finalists, including a fusion of the best models, answer conversations the search never saw. The winner is picked on these. 6. **Customers.** The finalists talk to a simulated customer who has the goal of a held-out conversation. When the search reaches its budget it stops and keeps its results. **Add $25** continues it from where it stopped. **Stop** ends it and keeps its results. If a Braintrust read fails while the search loads its cases, the worker tries again 3 more times — after 2 s, 10 s and 30 s — before it stops the search. A stopped search says why; fix the cause (reconnect the key, say) and add budget to continue. ## Read the results The chart plots each candidate's cost against its score, with the best-value frontier as a line. The table below it: | Column | Meaning | |---|---| | Model | The model, or `tm-mix-N` for a mix. Open a row to see what the mix does. | | Provider | The host that served the model, with its precision when the host states it. | | Score | The mean scorer score, in percent. | | $ / 1k req | The measured cost of 1,000 requests like your cases. | | p50 | The median time to a full answer. | | TTFT | The median time to the first token. | | Out tok | Output tokens per answer. | | Status | Which stage scored the row: `pilot`, `dev` or held-out. | The ★ row is the winner: the cheapest finalist that costs less than the baseline, is not worse than the baseline on a paired test over the same held-out cases, and scores within 2 points of it. Ties on cost go to the lower median latency. Rows marked `dev` or `pilot` were scored on fewer conversations and cannot win. If no row qualifies, there is no star. Open a row to see the cases where it did worse than the baseline, side by side, and its score per tag. ## Deploy the result **Deploy** makes a result callable as `tm/-v1`. The name is lowercase letters, digits and dashes, 40 at most: ```bash curl https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_KEY" \ -H "content-type: application/json" \ -d '{"model": "tm/billing-agent-v1", "messages": [{"role": "user", "content": "Hi"}]}' ``` - Only keys of your organization can call it. - Each model call it makes is billed like a direct call to that model. The response's `usage.cost` is their sum. - Send it on `POST /v1/chat/completions`. With `"stream": true` it returns the finished answer as one chunk, then usage. - Deploying again gives a new version (`-v2`). Earlier versions keep working. [Endpoints](https://app.routerplus.com/endpoints) lists everything your organization deployed, with the search it came from. **Try it** sends one held-out case of that search to the endpoint and bills your credits like any call. ## What is stored A search stores its cases and the answers candidates gave, for your organization only, until you delete the search. Deleting a search deletes them. A deployed model id keeps working after its search is deleted. --- Source: https://app.routerplus.com/docs/authentication.md # Authentication Every gateway request authenticates with a **virtual key** — one key that works on both wire surfaces, spends one org balance, and never touches a provider credential of yours. ## Virtual keys | Property | Value | |---|---| | Format | `tm_vk_` followed by 48 hex characters (24 random bytes) | | Shown | exactly once, at creation | | Stored | only the SHA-256 hash, plus the first 12 characters so the console can label it | | Spends | the org's shared prepaid balance | | Limits | per-key requests-per-minute; optional per-key and org-wide monthly spend caps | > [!WARNING] > The raw key is displayed once and cannot be recovered — only its hash is > stored. If you lose a key, disable it at [https://app.routerplus.com/console/keys](https://app.routerplus.com/console/keys) and create a new one. ## Both header forms are accepted The gateway reads your key from either standard header, on **both** surfaces: | Header | Convention | Typically sent by | |---|---|---| | `Authorization: Bearer tm_vk_...` | OpenAI | OpenAI SDKs; Claude Code via `ANTHROPIC_AUTH_TOKEN` | | `x-api-key: tm_vk_...` | Anthropic | Anthropic SDKs via `api_key` | You do not have to match the header to the surface — `x-api-key` works on `POST /v1/chat/completions` and `Authorization: Bearer` works on `POST /v1/messages`. If both headers are present, `Authorization: Bearer` wins. ```bash # OpenAI surface, bearer auth curl -s https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"hi"}]}' # Anthropic surface, x-api-key auth curl -s https://api.routerplus.com/v1/messages \ -H "x-api-key: $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"claude-haiku-4-5","max_tokens":50,"messages":[{"role":"user","content":"hi"}]}' ``` > [!NOTE] > `anthropic-version` is **not required** on `/v1/messages` — the gateway pins > its own version (`2023-06-01`) when it dispatches upstream, and forwards your > `anthropic-beta` header when the serving provider speaks the Anthropic > dialect — filtered to an allowlist (`prompt-caching`, `token-efficient-tools`, > `fine-grained-tool-streaming`, `interleaved-thinking`, `claude-code`, `oauth` > and `computer-use` betas); any other value is dropped, and the drop is named > in `x-tm-dropped-params`. The one place `anthropic-version` changes behavior: sending it on > `GET /v1/models` returns the Anthropic-shaped model list instead of the > OpenAI list shape. Anthropic SDKs send it automatically; that is fine. In the SDKs, the key goes exactly where the provider's own key would: ```python import os from openai import OpenAI import anthropic openai_client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) anthropic_client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"]) # anthropic.Anthropic(auth_token=...) also works — that sends Authorization: Bearer, # which the gateway accepts too. ``` ```typescript import OpenAI from "openai"; import Anthropic from "@anthropic-ai/sdk"; const openai = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY }); const anthropic = new Anthropic({ baseURL: "https://api.routerplus.com", apiKey: process.env.TM_API_KEY }); ``` ## Getting a key ### In the browser Open [Sign up](https://app.routerplus.com/signup) to create an account, or [Sign in](https://app.routerplus.com/login) if you already have one. Public signup and sign-in use Clerk. Complete the authentication and email-verification steps shown there. Use the same verified email as your existing marketplace account to keep its organization, keys and balance; switching sign-in methods does not create another trial. If you have an older marketplace account but no Clerk sign-in yet, use **Sign up** with that same email to connect it to your existing account. For a new account, the first verified sign-in mints your first API key. Signup and email verification do not grant free credit. Add paid credits in [Billing](https://app.routerplus.com/console/billing) before making model calls. Save the key when it is shown after sign-in; the raw value is shown only once. At [API keys](https://app.routerplus.com/console/keys) you can create additional named keys, each likewise shown once. Browser sign-in does not change how your applications authenticate with existing `tm_vk_` keys. Public programmatic signup and the local magic-link submission route have been removed: `POST /v1/signup` and `POST /login` return HTTP `404`. The `?magic=1` query parameter does not provide an alternate signup or sign-in form. Existing accounts do not need to be recreated; sign in through Clerk using their verified email. A marketplace session lasts up to 7 days and remains subject to the linked Clerk session. To end it sooner, use **Sign out** in the account menu (the avatar in the app bar) or at the bottom of the console rail. Sign out ends this browser's session on the server. Other devices stay signed in. ## Key rate limits Each key has a request limit over a rolling 60-second window, enforced per key and shared across gateway instances: | Key origin | Requests per minute | |---|---| | First key shown after verified sign-in (trial) | 20 | | Created in the console, or with the identity API | 300 | | The console's managed Playground key | 60 | Until your organization has bought credit, its customer keys and the organization as a whole run at the trial rate, 20 requests per minute, whatever a key's own limit says; the first purchase lifts it within seconds. Requests served only by your own provider keys (BYOK) and the managed Playground key are not held to it. Exceeding it returns `429` with `error_type: "rate_limit"` and a `retry-after` header. Honor the header and retry; don't hammer. Your organization also has token-per-minute and concurrency limits; see [Rate limits & spend caps](/docs/limits). ## Disabling a key At [https://app.routerplus.com/console/keys](https://app.routerplus.com/console/keys), each key row has a **Disable** action. Disabling is permanent — there is no re-enable in v1; create a new key instead. The gateway caches key rows briefly, so a disabled key can keep working for up to about five seconds before every request returns 401. You can also cap a key without killing it: per-key **monthly spend caps** are set on the same page (**Set cap**), and the org-wide cap at [https://app.routerplus.com/console/billing](https://app.routerplus.com/console/billing). A breached cap returns `429 insufficient_quota` with the exact UTC reset time in both the message and the `x-tm-cap-reset` header. ## What a 401 looks like A missing, invalid, or disabled key gets `401` with `error_type: "auth"`. The body is shaped for whichever SDK family is calling, so your SDK's native error class fires; the stable `error_type` rides alongside. Headers always include `x-request-id`, `x-tm-error-code: auth` and `x-tm-error-origin: authorization`. OpenAI surface (`/v1/chat/completions` and every non-Anthropic path): ```json { "error": { "code": "auth", "message": "missing or invalid API key", "type": "invalid_request_error", "metadata": { "error_type": "auth", "origin": "authorization", "retryable": false } }, "request_id": "..." } ``` Anthropic surface (`/v1/messages`): ```json { "type": "error", "error": { "type": "authentication_error", "message": "missing or invalid API key", "error_type": "auth", "metadata": { "origin": "authorization", "retryable": false } }, "request_id": "..." } ``` Both SDK families raise their own `AuthenticationError` for these. The message is deliberately the same for every 401 cause — check, in order: the header form, the key value, and whether the key was disabled in the console. ## Key hygiene - Keep keys in environment variables or a secret store; never commit them and never put them in URLs — the gateway only reads the two headers above. - Use one key per app or agent, so the console's per-key spend and caps tell you who spent what — and so disabling one key kills one integration, not all of them. - The console also mints keys of its own: one **Playground** key per organization, and one per Optimize search. They show on the keys page with a **managed** tag, are never displayed, and bill your credits like any key. Next: [Quickstart](/docs/quickstart) for browser signup and the first API calls, or [Errors](/docs/errors) for the complete `error_type` table. --- Source: https://app.routerplus.com/docs/byok.md # Bring your own provider key Your organization can route inference through its own OpenAI, Anthropic or Azure API keys, or an AWS Bedrock workload role. You still call the gateway with a marketplace key; the provider credential stays stored with the connection, encrypted, and the provider bills you directly. The marketplace charges nothing for BYOK inference. Set it up on [**Console → Provider connections**](https://app.routerplus.com/console/byok) or with the API below. ## Set up in the console 1. Sign in, open [`/console/byok`](https://app.routerplus.com/console/byok), and select your organization. 2. Click **Add connection**. Name it, choose a provider, enter its credential and add exact model IDs. Azure also needs a resource and a deployment mapping. Bedrock is selectable only when your organization has an operator-approved role. The provider and model list are fixed after saving; create a new connection if they change. 3. Click **Save and validate**. If validation fails, the pending connection remains available; rotate its credential or retry validation instead of creating a duplicate. 4. Choose an existing provider account's limits, or enter RPM, TPM, concurrency and headroom for a new account. Credentials for the same account should share the same pool or parent. A connection with no pool cannot serve requests. 5. Create a new platform key or select an existing one. A new BYOK key is created and bound in one transaction before it is returned. The page warns before it replaces an existing key's connection or key-level policy; an organization or workspace policy still applies. 6. Copy the new platform key once and use the generated request example. The optional **Send test request** sends "Reply OK" with at most 8 output tokens; it uses provider quota and may incur a small provider charge. The connection cards support validation, shared-account limit edits, rotation, disable and permanent deletion. Viewers can inspect metadata; owners and admins make changes. Provider secrets are never returned. A disabled connection needs rotation and successful validation to become active again. Deleted connections cannot be restored. Workspaces and principals provisioned through the [identity API](/docs/identity) can be selected when issuing a key. The console preserves their limits and restrictions. ## API setup 1. Sign in to the console and create a virtual key. No marketplace credit is required for BYOK inference. 2. Send `POST /api/connections` on the **site origin**, with JSON: ```json {"profile":"openai","models":["your-exact-provider-model-id"],"api_key":"YOUR_PROVIDER_KEY"} ``` Basic profiles are `openai` and `anthropic`; [Azure and policy setup](/docs/routing-policies) adds an `azure` profile with mandatory deployment mapping, and [Bedrock](/docs/bedrock) a `bedrock` profile whose `workload_role` replaces `api_key`. Supply 1–100 exact model IDs, without wildcards. An optional `name` labels the connection. The response is `201` with `connection_id` and `status: "pending"`; it never returns the key. 3. Send `POST /api/connections/{connection_id}/validate` with `{}`. A successful result has `valid: true`, `status: "active"`. For API-key profiles, validation performs one authenticated model-list GET, with a ten-second timeout and no generated tokens. It checks credential access to that endpoint; it does not verify access to every declared model or reserve capacity. 4. Send `POST /api/connections/{connection_id}/bind` with `{"key_id":"YOUR_VIRTUAL_KEY_UUID"}`. This binds an existing, enabled virtual key in the same org to an active connection. Allow five seconds for a previously used house key to switch, or use a fresh virtual key. 5. [Declare a quota pool](/docs/admission) for the provider account and assign the connection to it. A route with no declared pool is refused with a 503. Credentials for the same account must share the appropriate pool or parent. 6. Use the virtual key with `/v1/chat/completions` or `/v1/messages`. Both buyer formats and streaming work with either profile, subject to the existing wire compatibility limits. Management requires the `tm_s` session cookie. Mutations also require `Content-Type: application/json` and an `Origin` header equal to `https://app.routerplus.com`. Same-origin browser `fetch` supplies the cookie and Origin automatically. A CLI client must supply both securely; avoid putting cookies or provider keys in shell history. An inference bearer key cannot manage connections. Viewers can read; owners and admins can write. ## Management operations | Method and site path | JSON body | Effect | |---|---|---| | `GET /api/connections` | none | Newest 100 org-owned connections, including tombstones | | `GET /api/connections/{id}` | none | One connection's metadata and fully masked label | | `POST /api/connections` | `profile`, `models`, `api_key` or `workload_role`, optional `name`, `endpoint_config` | Create a pending connection | | `POST /api/connections/{id}/validate` | `{}` | Check credential, activate on success; failed validation leaves pending | | `POST /api/connections/{id}/rotate` | `api_key` or `workload_role` | Revoke old versions and create a pending version; validate it before use | | `POST /api/connections/{id}/disable` | `{}` | Disable connection and revoke its credential versions | | `POST /api/connections/{id}/bind` | `key_id` | Bind/rebind an org-owned virtual key to an active connection | | `DELETE /api/connections/{id}` | `{}` | Revoke and tombstone; retain attribution and bindings | Disabled connections can be restored by rotating and validating a new credential. Deleted connections cannot be restored. Rotation temporarily stops new inference until validation succeeds. Profile and allowed-model changes require a new connection and an explicit rebind. There is no implicit unbind-to-house operation. An explicit house-only policy is available through the [policy API](/docs/routing-policies). No secret suffix, ciphertext, envelope, or upstream validation body is exposed in management responses. ## Inference behavior A bound key can use only its connection's exact models. It never silently falls back to marketplace supply. `/v1/models` lists that connection's models; Anthropic token counting uses the same connection and is available only for an Anthropic profile. Unavailable, disabled, or deleted bindings produce `404 model_unavailable`; database/decryption failures produce `503 gateway_error`. An invalid virtual key still produces `401 auth`. BYOK inference is metered and durably audited. Marketplace charges and marketplace provider payables are zero: the attempt's `billing_source` is `byok`, its `cost_usd` is `0`, and its inference cost is recorded as unknown (`null`), since the marketplace does not price your account. The provider still bills your account. Marketplace spend caps do not limit that external bill; virtual-key RPM enforcement and your organization's token and concurrency limits still apply, including to token counting. A policy with house fallback is the one case a BYOK key spends credit: the house attempt reserves and is charged like any house call. Image generation works through an `openai` connection: list the exact image model ids you will send (matching is literal), and the render is billed by OpenAI to your own account, so `usage.cost` is `0`. See [POST /v1/images/generations](/docs/api-images). New calls reflect credential disable/rotation within five seconds; existing upstream calls and streams may finish on their original version. When the configuration database is unreachable, previously verified authority may serve for a bounded time; see [limits and outage behavior](/docs/admission). OpenAI, Anthropic and configured Azure v1 public destinations are supported; redirects are refused. Bedrock uses the separately documented workload-role profile. No arbitrary endpoint or automatic private-link, residency or retention guarantee is included. [Routing policies](/docs/routing-policies) add explicit ordering, filtering and funding controls. Request-level credentials or destinations (`api_key`, `base_url` and the like) return 400. Provider HTTP error bodies are never relayed, so an echoed credential cannot leak. Validation and token-count calls are not inference ledger attempts; validation has a content-free management audit event. Shared limits and pilot fairness are described in the [admission guide](/docs/admission). ## Related Use [workspaces and principals](identity.md) for scoped keys and trusted end-user limits. The [Bedrock workload-role profile](bedrock.md) adds native regional Claude inference; its write-only `workload_role` replaces `api_key`. The connection lifecycle, quota-pool requirement and free BYOK economics are the same for every profile. --- Source: https://app.routerplus.com/docs/routing-policies.md # Routing policies and Azure connections A saved routing policy says which connections and marketplace deployments a key may use, in what order, and who pays. Policies apply to orgs, workspaces and keys. ## Create and bind a policy Use the site origin and the same session cookie, JSON content type, and same-origin mutation protection as [BYOK connection management](/docs/byok). Inference keys cannot administer policies; viewers can read them. `POST /api/routing-policies` creates an immutable policy revision: ```json { "models": ["your-exact-model-id"], "connections": ["FIRST_CONNECTION_UUID", "SECOND_CONNECTION_UUID"], "funding": "byok_only", "allow_fallbacks": true, "max_attempts": 4, "timeout_ms": 450000 } ``` The response is `201` with `{"policy_id":"UUID"}`. Connections must belong to your org. Bind an existing key with `POST /api/routing-policies/{id}/bind`, body `{"key_id":"KEY_UUID"}`. Bind an org policy with the same endpoint and `{"scope":"org"}`. Org policies apply to all keys in the org; key policies can narrow them. A key without its own policy or single connection binding uses the org policy. A legacy single-connection binding remains an additional restriction when an org policy exists. Policy binding replaces a key's prior single-connection binding; binding a single connection replaces its key policy. `GET /api/routing-policies` lists the newest 100 revisions. GET by policy UUID returns one. Updates create a new revision and explicitly rebind it. There is no delete or implicit unbind-to-house operation. House-only cached decisions may take up to five seconds to observe a new binding. Existing dispatched calls may finish on their pinned version. ## Funding, ordering and limits | Field | Behavior | |---|---| | `models` | 1–100 exact IDs; all applicable policies must allow the requested model | | `connections` | Up to eight ordered org-owned BYOK connection UUIDs | | `house_deployments` | Up to eight explicit marketplace deployment IDs | | `funding: "byok_only"` | Requires connections, forbids house deployments; default | | `funding: "house_only"` | Requires house deployments, forbids connections | | `funding: "byok_with_house_fallback"` | Requires both; house remains the final phase | | `allow_fallbacks` | Defaults true; false limits selection to the first eligible route before health filtering | | `max_attempts` | Integer 1–8; default four | | `timeout_ms` | Generation/admission execution deadline, 1,000–600,000 ms; default 450,000 | | `regions`, `zdr`, `data_collection` | Hard restrictions against operator-supplied evidence; missing evidence fails closed | Under a policy, automatic fallback stays within the first selected **provider label**. OpenAI and Azure can both be authorized and explicitly selected, but a failed OpenAI call does not automatically switch to Azure. Exact-model equivalence remains a separate decision; `attest_equivalence:true` is rejected. Azure deployment mappings are customer declarations. The effective restrictions are the intersection of org policy, workspace policy, key policy/binding, the server catalog, and request restrictions. Empty intersections fail closed. Lower limits win. Ordering uses request connection preference, then provider preference, then saved order and a stable deployment ID tie-breaker. Marketplace supply remains after all BYOK candidates for mixed funding. An unavailable primary does not authorize a fallback when fallbacks are disabled. No retry crosses the existing streaming commit boundary or a content-policy denial. The execution deadline covers admission and generation after planning, including active streams. Request upload, bounded policy/price reads, and mandatory journal settlement/drain are separate phases. Timing out never abandons required durable settlement. A late SQL reservation is released and never dispatched. ## Request controls and explanations The inference `provider` object accepts: ```json { "only": ["openai"], "ignore": [], "order": ["openai"], "connections": ["CONNECTION_UUID"], "connection_order": ["CONNECTION_UUID"], "funding": ["byok"], "allow_fallbacks": false, "require_parameters": true, "upstream": ["deepinfra/bf16"] } ``` `only`, `ignore` and `order` name route labels: house `anthropic`, `openai` and `openrouter`, and a BYOK connection's profile (`openai`, `anthropic`, `azure`, `bedrock`). A route label keeps that meaning even when its route does not serve the model. Any other label is an OpenRouter host tag. When the `openrouter` route runs, the gateway sends the host tags to OpenRouter as its own `order`, `only` and `ignore`, with your `allow_fallbacks` (default true). On `POST /v1/decisions` the labels are the decision providers: `typesafe` and `openrouter-decisions` for Jev, `openrouter-decisions-free` for Mercury Decide, `bespokelabs` for Bespoke Nimble v3, `workers-ai` for Clef and Clef-flash, `routerplus` for Decider 2B and Kev 4B, `perplexity-decisions` for Perplexity Decider v1 27B, `levanto` for Sage and `fastino` for GLiNER-2.5-Decide. Clef's label is `workers-ai`, not `cloudflare`: `cloudflare` stays an OpenRouter host tag (OpenRouter's Cloudflare host runs some chat models), and on `POST /v1/decisions` it names no provider. In the same way, Perplexity Decider's label is `perplexity-decisions`, not `perplexity`, which stays OpenRouter's host tag for Perplexity's chat models. Route labels win over host tags with the same spelling: `only: ["anthropic"]` means the Anthropic route. A host tag in `only` keeps the `openrouter` route open. `upstream` names hosts behind an aggregator deployment (OpenRouter provider tags), up to eight, tried in that order with no fallback to other hosts; with `upstream`, the host tags from `order`, `only` and `ignore` do not go. The tags are OpenRouter provider slugs; the hosts that run a model are its `served_by` entries in [`/api/models.json`](https://app.routerplus.com/api/models.json). To keep a session's prompt cache on one host, see the [Agent integration guide](agent-integration.md). `require_parameters: true` refuses a deployment that would drop one of your request fields instead of dispatching without it. `regions`, `zdr` and `data_collection` are accepted here too as hard restrictions. These preferences never grant connections, models, providers, or funds that saved authority does not allow. Price/latency/throughput sorting, max-price controls, raw request keys and arbitrary destinations are unsupported and return 400. Customers cannot write the evidence used to satisfy region and privacy restrictions, and a restriction is not a residency or ZDR guarantee. `POST /v1/route` on the gateway accepts `{"model":"ID","provider":{...}}` using the inference key. It returns the plan, catalog and policy IDs, the ordered candidates with their funding and pinned house price IDs, the exclusions with a reason each, the OpenRouter host tags it would send as `openrouter_provider` (`null` when there are none), and the signal time and source, with `advisory: true`. It performs no provider request, decryption or reservation. Results are advisory; health and authorization can change before actual dispatch. `GET /v1/models` uses the same restrictions. An optional URL-encoded JSON `provider` query parameter narrows the listing, and `output_modalities=text,image,video,decisions` filters by kind. Token counting uses the same plan but requires an Anthropic transport; it sends one count request and does not perform inference fallback. `GET /v1/generation?id=...` includes each attempt's durable `route_context` and fallback cause. Policy inference returns `x-tm-route-plan-id` for correlation. Free BYOK still requires no marketplace credit. House fallback reserves credit only when reached and charges under the existing pricing rules. Earlier house holds remain until the request finishes, so fallback admission is conservative while settlements are buffered. Each attempt pins its own funding/credential identity. BYOK external cost stays unknown; house prices come from the existing model-level registry and are frozen before dispatch. ## Azure/Foundry API-key connections Create a connection through `POST /api/connections`: ```json { "profile": "azure", "models": ["your-public-model-id"], "api_key": "YOUR_RESOURCE_API_KEY", "endpoint_config": { "resource": "your-resource-name", "host": "foundry", "api_version": "v1", "deployments": {"your-public-model-id": "your-deployment-name"} } } ``` `host` is `openai` (`{resource}.openai.azure.com`) or `foundry` (`{resource}.services.ai.azure.com`); the server constructs the public hostname, and customers cannot submit URLs. Every declared model needs a deployment mapping. Endpoint config is immutable; changes require a new connection. Validate, then bind it directly or include it in a policy. Rotation and revocation follow the connection lifecycle. The adapter uses `/openai/v1/chat/completions?api-version=v1`, the deployment name in the model field, and the native `api-key` header. Validation calls the model-list endpoint with no generation. These shapes follow Microsoft's [chat reference](https://learn.microsoft.com/en-us/rest/api/microsoft-foundry/azureopenai/chat), [models reference](https://learn.microsoft.com/en-us/rest/api/microsoft-foundry/azureopenai/models), and [v1 lifecycle guide](https://learn.microsoft.com/en-us/azure/foundry/openai/api-version-lifecycle). Only API-key authentication and public v1 endpoints are implemented. Entra/workload identity, legacy dated APIs, preview versions, sovereign/private endpoints, and Azure capacity/residency qualification are not included. Validation confirms endpoint authentication, not deployment entitlement, exact-model equivalence, quota, or capacity. Shared admission requires declared [quota pools and gateway limits](/docs/admission). Quota availability may skip a candidate inside this frozen authority; it never grants new providers or marketplace funding. Model listings/explanations do not reserve future capacity. Workspace bindings use the [identity API](identity.md). Durable attempt evidence stores at most 32 exclusion details and `excludedTotal` for the full count; the catalog hash and full route explanation preserve context without letting a large catalog exceed journal bounds. --- Source: https://app.routerplus.com/docs/identity.md # Workspaces, members and principals Every inference key belongs to one org, workspace and principal. These bindings are immutable: mint a replacement key to change identity. Multiple keys for one principal share its RPM, TPM and concurrency limits. Request `user`, `metadata`, `x-user-id`, `x-tm-principal-id` and `x-tm-workspace-id` values do not authenticate a principal. Your trusted backend provisions an end-user principal and retains that principal's key; do not let an end user choose a different backend key. There is no per-request delegation, SSO or SCIM. Existing keys and keys created through browser onboarding or the console use an org's **Default** workspace and shared `_default` service principal. This preserves existing access. It does not retroactively identify individual users of a shared legacy key. ## Management authentication Use a verified console session cookie. Inference bearer keys have no management rights. `x-tm-org-id: ` selects an org in which the session email must be an active member; it never grants access. Without the header, the caller's own org is preferred, then their oldest active membership. Query `/api/identity/organizations` to discover memberships. The console shows your default org; [`/console/byok`](https://app.routerplus.com/console/byok) and the playground take `?org=` to switch, and the other scoped identities are managed with the API below. POSTs require `Origin: https://app.routerplus.com` (origin only) and `Content-Type: application/json`. Responses are not cacheable. No API sends an invitation or email when adding membership. Membership and member-principal emails use the same canonicalization as marketplace account emails (lowercase, plus-tag removal, and Gmail dot/domain normalization). | Role | Read org configuration | Manage workspaces, principals, keys, connections, policies, limits | Manage members | |---|---|---|---| | owner | Yes | Yes | Yes | | admin | Yes | Yes | No | | viewer | Yes | No | No | Roles apply across the whole org. There are no workspace-specific human roles in this version. The initial owner cannot be disabled, demoted or replaced through this API. Removing a membership removes that session's authority in this org without invalidating its memberships elsewhere. ## Identity endpoints All collection GETs return at most 500 entries; use returned UUIDs for subsequent writes. A create returns `201`; an update, and a member write, `200`. | Request | JSON body / result | |---|---| | `GET /api/identity/organizations` | `{organizations:[{org_id,name,role}]}` | | `GET /api/identity/members` | `{members:[{email,role,disabled}]}` | | `POST /api/identity/members` | `{email,role:"admin"\|"viewer",disabled?:boolean}`; owner only, creates or updates | | `GET /api/identity/workspaces` | Workspace metadata and policy pointers | | `POST /api/identity/workspaces` | `{name}` → `{workspace_id}` | | `POST /api/identity/workspaces/` | Any of `{name,disabled,route_policy_id}`; `null` clears its policy | | `GET /api/identity/principals` | Principal metadata | | `POST /api/identity/principals` | `{workspace_id,kind:"service"\|"end_user"\|"member",subject,member_email?}` → `{principal_id}` | | `POST /api/identity/principals/` | `{disabled:boolean}` | | `GET /api/identity/keys` | Metadata only, including workspace/principal bindings | | `POST /api/identity/keys` | `{principal_id,name,connection_id?}` → `{key_id,key}`; raw key shown once; 300 RPM | | `POST /api/identity/keys/` | `{disabled:boolean}` | `subject` is an opaque stable backend identifier, unique within a workspace, limited to 160 characters. Leading `_` is reserved. Use `kind:member` with an existing member's `member_email` to make membership revocation also stop their inference keys. An `end_user` or `service` principal is independent of human membership; disable it directly to revoke it. Supplying `connection_id` while issuing a key checks that the connection is active and owned by the same org, then creates the key with that binding in the same transaction. The BYOK console uses this path so an interrupted setup never reveals a temporarily house-funded key. Omitting it preserves the original key-creation behavior for non-BYOK callers. Disabling a workspace denies all its keys. Disabling a principal denies all its keys. Disabling a member denies their control access and their `member` principals. New dispatches observe the existing five-second revocation maximum; a previously dispatched stream can finish. A successful mutation response waits for publication of its Redis fence. If it returns 503, inspect state before retrying: SQL may have committed before fence publication. ## Routes, limits and evidence Effective authority is the intersection of **org → workspace → key → request** restrictions. Attach an immutable routing policy with `POST /api/identity/workspaces/` and `{route_policy_id:"..."}`. This cannot grant a route denied by an org or key policy. The `/api/admission-limits` API accepts `scope:"workspace"` and `scope:"principal"`, with the corresponding UUID in `subject`. Defaults for each are 300 RPM, 1,000,000 TPM, eight concurrent requests. All org/key/model/platform/provider-pool limits also apply. These rate limits do not create separate workspace/principal monetary balances or spend caps; money remains org/key scoped. Configure `outage_mode:"closed"` when contractual limits require refusing work during counter failure. `GET /v1/usage` and `/v1/generation?id=...` expose attempts belonging to the authenticated principal and workspace, including its other keys. Unknown or foreign generations return 404. Org balance fields are `null` for explicitly scoped principals; the default service principal retains those fields. The console provides org-wide accounting to authorized members. Attempt route evidence records `workspaceId`, `principalId` and up to three policy revisions. --- Source: https://app.routerplus.com/docs/bedrock.md # Bedrock workload identity The `bedrock` connection profile uses AWS SigV4 and temporary AssumeRole credentials. It supports **regional Anthropic Claude Messages through InvokeModel and InvokeModelWithResponseStream**, exposed through both existing buyer surfaces. It is not a universal Converse adapter. `count_tokens` is not supported for this profile. Only direct `anthropic.claude-...-vN:N` model IDs are accepted. Cross-region/global inference profiles, profile ARNs, provisioned-throughput ARNs and arbitrary endpoints are rejected. The public model-to-Bedrock model mapping and AWS region are immutable connection config. Availability and account model access must be checked for the exact chosen region. ## Operator setup The gateway's AWS credentials come from its instance role, through IMDSv2; the application does not use an ambient AWS static-key credential chain. Grant this role `sts:AssumeRole` on the customer's exact role ARN. The customer role trust policy must name the source role and require `sts:ExternalId` equal to `tm:`. Its permissions should allow `bedrock:InvokeModel` and `bedrock:InvokeModelWithResponseStream` only on approved regional foundation-model ARNs. Both processes require an operator-controlled grant map, e.g.: ```sh TM_BEDROCK_ROLE_GRANTS='{"ORG_UUID":["arn:aws:iam::123456789012:role/tm-customer"]}' ``` An org cannot register or use another org's role without an explicit operator grant. Updating the map requires restarting the affected processes. It does not replace AWS IAM authorization. Keep trust grants least-privilege and use a distinct external ID per org. The console offers the Bedrock profile only to an organization with a grant. ## Create and validate Using the same session/Origin headers as other connection management calls: ```json { "profile": "bedrock", "models": ["your-public-model-id"], "endpoint_config": { "region": "us-east-1", "deployments": { "your-public-model-id": "anthropic.claude-3-haiku-20240307-v1:0" } }, "workload_role": { "role_arn": "arn:aws:iam::123456789012:role/tm-customer" } } ``` POST this to `/api/connections`; replace the example model with a regional model your account can invoke. The role reference is write-only and envelope-encrypted using the existing environment wrapping-key ring. Do not supply `api_key`, AWS access keys, an external ID override, a bearer token or a request-time credential. `POST /api/connections//validate` checks the AssumeRole path without generating a billable completion. An active result proves role assumption, **not model entitlement, capacity, private connectivity, processing residency or ZDR**. Declare and bind a `provider:"bedrock"` quota pool, then bind the connection or policy to a virtual key. Run a real minimal inference to qualify model permissions and streaming. Rotation uses `/rotate` with `{workload_role:{role_arn:"..."}}`, revokes the old connection version and returns it to pending. It preserves account-pool identity. Disabling or deleting the connection uses the existing fenced revocation path. ## Runtime behavior STS uses the configured regional endpoint, 15-minute sessions and an org-bound external ID. Temporary credentials live only in bounded process memory (256 sessions, 64 concurrent refreshes), refresh at least a minute before expiry and never enter SQL, Redis, journals or telemetry. Both AWS clients use one attempt; the gateway remains the retry authority. The HTTP/1.1 handler is set explicitly. SDK-internal retries and endpoint-environment overrides are disabled. The durable intent precedes role resolution and inference. The existing final authority and admission checks still precede the signed inference call. A credential failure before inference is accounted as never dispatched. Inference transport uncertainty retains the normal conservative usage reservation. AWS binary events are decoded into the existing Anthropic stream pipeline; once output commits, errors are terminal, with no silent retry. AWS exception text is not relayed because it can echo sensitive input. Free BYOK economics are unchanged: platform debit and provider payable are zero; external inference cost remains unknown unless independently priced. A regional endpoint is not a ZDR attestation. Privacy/residency route filters still require operator evidence. References: [AWS Claude request format](https://docs.aws.amazon.com/bedrock/latest/userguide/model-parameters-anthropic-claude-messages-request-response.html), [stream operation](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_InvokeModelWithResponseStream.html), [Bedrock PrivateLink](https://docs.aws.amazon.com/bedrock/latest/userguide/vpc-interface-endpoints.html). --- Source: https://app.routerplus.com/docs/admission.md # Rate limits, capacity and outage behavior Admission is shared across gateway instances: every limit on this page is counted in one store, so two instances never admit the same minute twice. This page is the full picture — scopes, provider pools, fairness, and what happens when a store is down. The everyday view (per-key RPM, the org defaults, spend caps) is on [Rate limits & spend caps](/docs/limits). BYOK connections and routing policies, which pools attach to, are on [BYOK setup](/docs/byok) and [routing policies](/docs/routing-policies). ## Manage limits Sign in and open [`/console/limits`](https://app.routerplus.com/console/limits) to view or set gateway limits, create provider pools and assign connections. Pool updates replace the supplied limits and parent; enter the complete desired configuration. Gateway limits overlap: the strictest applicable platform, org, workspace, principal, key, model and provider-account constraint applies. Management APIs use the site origin (`https://app.routerplus.com`), session cookie, JSON content type and same-origin mutation protection. Inference bearer keys cannot administer them. Owners and admins can write; viewers can read. | Endpoint | Purpose | |---|---| | `GET /api/admission-limits` | List your org's configured limits, every scope | | `POST /api/admission-limits` | Replace one scope's limits | | `GET /api/quota-pools` | List the newest 100 org-owned provider pools | | `GET /api/quota-pools/{id}` | Read one pool | | `POST /api/quota-pools` | Create a pool | | `POST /api/quota-pools/{id}` | Replace its limits and parent | | `POST /api/quota-pools/{id}/bind` | Assign an owned connection to that pool | Example gateway limit body: ```json {"scope":"key","subject":"VIRTUAL_KEY_UUID","rpm":100,"tpm":200000,"concurrency":4,"outage_mode":"bounded_open"} ``` `scope` is `org`, `workspace`, `principal`, `key` or `model`. For org scope, subject is `org`; for workspace, principal and key scopes it is the UUID; for model scope it is the exact model ID. Null/omitted numeric values inherit defaults or parent restrictions. RPM accepts 1–1,000,000, TPM 1–1,000,000,000, concurrency 1–10,000. A key's original RPM ceiling still applies, so configuring a larger key limit does not raise that ceiling. Use `outage_mode:"closed"` for explicit refusal during counter-store failure; default is `bounded_open`. A closed mode at any applicable scope wins. Default limits are 300 RPM, 1M TPM and 8 concurrent requests for an org, a workspace and a principal. Until an org has added credit, the org and each of its keys are held to 20 RPM (an operator override replaces this, as it replaces the defaults). Keys default to their RPM ceiling, 1M TPM and 8 concurrent requests. A model scope has no default: set it to add a limit for one model. The operator configures platform limits independently. Under a contract, the operator can also replace the defaults and the keys' RPM ceiling for one org (an override, with a per-second limit on the org if needed); limits you set still apply and can only lower them. A [dedicated endpoint's](/docs/dedicated-endpoints#limits) traffic keeps the endpoint's own limits. ## Declare provider account capacity ```json { "provider":"openai", "rpm":500, "tpm":1000000, "concurrency":20, "headroom_percent":90, "parent_pool_id":null } ``` Use `openai`, `anthropic`, `azure` or `bedrock` for customer pools. Supply finite positive values. `headroom_percent` is the usable percentage, from 1–100; default 90. Fractional limits round down, with a minimum of one unit. A pool parent must be owned by the same org and use the same provider. Up to four acyclic levels are supported, including the leaf. Bind with `{"connection_id":"CONNECTION_UUID"}` at `/api/quota-pools/{id}/bind`; the connection's profile must match the pool's provider. Keys for the same upstream account should share a pool, or child pools under a shared account parent. Rotating credentials does not reset capacity. Parent and child claims are checked together. The service cannot infer your real provider-account identity from a key; accurate pool declarations remain your responsibility. Marketplace house pools are operator-owned and mapped to catalog deployment IDs; customers cannot change them. A route with no declared pool is unavailable (`503`, `x-tm-limit-kind: capacity_unknown`), rather than implicitly unlimited. Provider usage outside this gateway is not observed. Provider 429s cool the affected leaf pool for the provider's `retry-after` (1 to 60 seconds), shared across workers; an unknown error scope is not promoted to a global provider outage. ## Counting and token estimates One admitted logical inference request consumes gateway RPM once. Every physical attempt, including an eligible fallback, separately consumes upstream RPM and token reservation. Requests already sent to a provider are not refunded as if they never happened. A request proven not dispatched may release its provider claim. Concurrency uses renewable ten-second leases so a crashed worker does not strand slots. Input reservation uses UTF-8 byte size plus a conservative media allowance (65,536 per image or document part). The base64 of an image sent inline is left out of the byte size: the image counts its allowance only. It is an estimate, not a vendor tokenizer guarantee. Requests enforce an output bound of 1–32,768 tokens, default 4,096, on the outgoing wire. Larger or conflicting output limits return 400. Completed, observed input and output usage replaces the estimate; cache and reasoning subsets are not counted twice. Partial prefill is not a final bill: cancelled, incomplete, unknown or estimated usage keeps the reservation until its rate window expires, except for proven never-dispatched work. The rolling-minute implementation includes a conservative partial second, so a charge can remain for up to 61 seconds. Anthropic token counting uses the same authorization and shared limits. It consumes gateway and provider RPM, reserves estimated input capacity and records a durable request audit before calling upstream. It does not create an inference debit or inference-attempt row. Its answer is `{"input_tokens": N}` and nothing else. ## Inspect limits and errors `GET /v1/limits?model=MODEL_ID` on the gateway returns effective gateway limits for the key, including inherited restrictions, the output bound and outage mode. Model is optional. Successful admission exposes `x-tm-remaining-rpm` and `x-tm-remaining-tpm`, the minimum remaining amount across claimed scopes at that moment. These are snapshots, not promises that capacity will still be available for a later request. On a dedicated endpoint with a [paced limit](/docs/dedicated-endpoints#the-paced-limit), `x-tm-queue-ms` says how long the request waited for its turn before admission. Failures retain SDK-compatible statuses and include metadata and CORS-readable headers: | Origin | Typical status | Meaning | |---|---:|---| | `gateway_admission` | 429 or 503 | Customer quota/spend limit, fairness headroom or worker capacity | | `upstream_quota` | 429; 503 if unconfigured | Declared account pool exhaustion/cooldown or unknown capacity | | `upstream` | Provider status | An actual provider response, including its 429 | | `gateway_infrastructure` | 503 | Counter/journal/recovery failure | | `authorization` | 401, 404 or 503 | Invalid/unavailable authority or changed/expired snapshot | Headers include `x-tm-error-origin`, `x-tm-limit-scope`, `x-tm-limit-kind` (`rpm`, `rps`, `rps_queue`, `tpm`, `concurrency`, `cooldown`, `spend`, `health`, `capacity_unknown`, …), optional `x-tm-limit-id`, and `retry-after`. Metadata includes retryability and, where known, retry or cap-reset time. A rolling-window retry time is advisory; it is not a capacity guarantee. Quota/admission rejections after authentication are recorded even without a physical attempt. The record is best effort: when it cannot be written, you still get the 429, not a 503. Query `GET /v1/generation?id=REQUEST_ID`: `admission_events` accompany attempts, and attempts include receipt context/error origin. Unauthenticated requests and HTTP-parser protections cannot be attributed to an authenticated org in this audit. ## Availability and fairness - **Healthy shared store:** limits coordinate across workers. Pilot fairness protects idle tenant headroom and caps each org at a share of a shared house pool's usable capacity — half by default, 75% on the decision-model pools (`typesafe`, `openrouter-decisions`, `openrouter-decisions-free`, `bespokelabs`, `workers-ai`, `levanto`, `fastino`, `routerplus`, `perplexity-decisions`) — with a one-unit minimum. It rejects excess work immediately; there is no waiting queue, except on a dedicated endpoint with a [paced limit](/docs/dedicated-endpoints#the-paced-limit), and no general no-starvation guarantee for arbitrary tenant counts. - **Redis unavailable:** the default bounded local allowance is at most five logical requests/minute, two concurrent local requests and 25,000 estimated TPM per scope per worker, for at most 30 seconds. A dedicated endpoint's traffic instead gets the endpoint's own limits divided among the gateways that serve it, for at most 10 minutes by default; a paced endpoint paces each gateway at its share of the rate. `x-tm-admission-mode: bounded_local` marks this path. Limits and balances can be approximate across workers during this interval. Strict mode refuses new work. Every provider call still requires a durable journal intent. - **Postgres unavailable:** previously verified snapshots may serve for a bounded time (five minutes by default; the operator can set up to one hour for an outage) while the live Redis authority fence still matches. A single failed read falls back to a snapshot at most five minutes old. Once reads have failed for five seconds with none succeeding, the gateway answers from the snapshot at once, up to 15 seconds after the last failure, instead of waiting on Postgres again. Cold, evicted, expired or changed authority fails closed. `x-tm-authorization-mode: bounded_snapshot` identifies requests that used this fallback. If both authorities are unavailable, cached authority is not used. - **Management during Redis failure:** mutations refuse rather than losing revocation coordination. A failed publication can leave a deliberate fence that needs operator recovery. Existing provider calls may finish on their pinned credential version. - **Redis process replacement/lost state:** shared counters require explicit recovery; an empty store is not treated as a full unused allowance. These are pilot controls. They do not establish provider capacity commitments, exact vendor TPM compliance, private connectivity, regional processing or ZDR. Under a routing policy, automatic fallback stays within one provider label (D10); see [routing policies](/docs/routing-policies). ## Workspace and principal scopes `workspace` and `principal` limits aggregate across the keys bound to that scope, inside the same atomic claim as org/key/model/platform/pool constraints. Bindings come from server-managed keys, never request headers. See [identity management](identity.md). --- Source: https://app.routerplus.com/docs/dedicated-endpoints.md # Dedicated endpoints A dedicated endpoint is provider capacity reserved for your organization on one model. It has its own limits, prices and failover rules. We set one up with you under a contract; there is no self-serve setup. You call it with the model id we give you, in your organization's own namespace: `/`, for example `acme/gemma-31b`. ## Calling an endpoint Use the endpoint's model id in place of a catalog model id. Everything else stays the same: both surfaces (`/v1/chat/completions` and `/v1/messages`), streaming or not, and every request field the model accepts. ```bash curl https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "acme/gemma-31b", "messages": [{"role": "user", "content": "Hello"}]}' ``` Only your organization's keys can call the endpoint. It can also be limited to named workspaces. For any other key the id does not exist: it gets the same 404 `model_unavailable` as an unknown model. On an open endpoint, the response names the model that ran, for example `"model": "google/gemma-4-31b-it"`, not the endpoint. A closed endpoint names the endpoint (see [Closed endpoints](#closed-endpoints)). Two headers tell you how the request was served: | Header | Value | |---|---| | `x-tm-dedicated-endpoint` | the endpoint's id, for example `acme/gemma-31b` | | `x-tm-route-role` | `primary` (your dedicated capacity) or `fallback` (the shared pool) | In [`/v1/generation`](/docs/api-usage), each attempt's `route_context.dedicated` records the endpoint id, the revision of its configuration in force, and the role. `GET /v1/models` lists your organization's active endpoints, with `owned_by` `routerplus`, on the base URL the endpoint is served from. The console's [Endpoints](https://app.routerplus.com/endpoints) page lists them first, before the models you deployed from Optimize: for each endpoint its id, model, base URL, status, provider, price per million tokens (input, cached input and output) and limits. ## What serves your request 1. **Your dedicated capacity first.** 2. **Then the shared pool, for the same model only.** The request moves on when the dedicated deployment fails before your answer starts, when its circuit is open, or when its capacity is full. It goes to public deployments of the same model that meet your endpoint's constraints. It never goes to a different model. 3. **The commit boundary holds.** Once the first token reaches you, the request stays with that provider. See [Routing & failover](/docs/routing). The constraints are part of your endpoint's configuration and apply to every fallback: | Constraint | A fallback host qualifies only if | |---|---| | precision | it serves one of the listed precisions, for example `bf16` | | zero data retention | it keeps no prompts | | data collection denied | it neither stores nor trains on requests | | maximum price | its list price per million tokens is at or below the limit | | regions | it runs in one of the listed countries | | hosts | it is one of the named hosts | | require parameters | it accepts every parameter the request sends | A host that has not stated a value for a constrained property does not qualify. A fallback that reaches several hosts through one route sends the constraints with the request, so only hosts that meet them serve it. A region constraint excludes such a route, because it makes no promise about where a host runs. An endpoint can also be set to `primary_only`. It then never falls back; when the dedicated deployment cannot serve, you get the error. For one request you can narrow the routes but never widen them. `"provider": {"allow_fallbacks": false}` keeps the request on your dedicated capacity. On open endpoints, `provider.only` and `provider.ignore` filter by provider name, as on any request. Closed endpoints select their providers and hosts through the contract configuration: `provider.only`, `provider.ignore`, `provider.order`, `provider.upstream` and `provider.require_parameters` return `400 invalid_request`, including empty selectors. Omit those fields on closed endpoints. `allow_fallbacks` remains supported. ## Closed endpoints An endpoint is open unless your contract says otherwise: like any request, it tells you which deployment and host of the shared pool served you, and your dedicated capacity shows as `routerplus`. A closed endpoint never names a deployment or a host. Every surface you can read names RouterPlus as the provider, whichever route served the request, and apart from the route role nothing you can read differs between the routes: | Surface | On a closed endpoint | |---|---| | `x-tm-provider` | `routerplus` | | `x-tm-served-by`, `x-tm-upstream-model`, `x-tm-upstream-status`, `x-tm-dropped-params`, `x-tm-affinity-wait-ms` | absent | | `x-tm-dedicated-endpoint`, `x-tm-route-role` | as on an open endpoint | | `x-tm-remaining-rpm`, `x-tm-remaining-tpm` | counted against the endpoint's own limits only | | Chat completion body | a fixed set of fields, below; `model` is the endpoint id | | Stream chunks | the same fields in each chunk; a provider's SSE comment lines are dropped; a provider's error in the stream ends it with the gateway's own error event, and nothing follows that event | | Error messages | they name `routerplus`, never a deployment; a provider's own error body is never relayed | | Error statuses | a provider's 401, 402, 403 or 404 is a retryable 503 `upstream_error`; a `retry-after` is the gateway's own, never a provider's | | Limits | a refusal because the capacity behind one route is full has `x-tm-limit-scope: endpoint` and `x-tm-limit-id: endpoint:` | | `provider.only`, `provider.ignore`, `provider.order`, `provider.upstream` | refused with 400 `invalid_request` | | [`/v1/generation`](/docs/api-usage) | `deployment` is `routerplus` and `model` the endpoint id; `route_context` keeps only `dedicated` (the endpoint, its revision and the role); `admission_context` is null, `dropped_params` is empty, and admission events carry no pool scope id | | [`/v1/usage`](/docs/api-usage) | `deployment` is `routerplus` and `model` the endpoint id | | Console: Overview, Usage, Logs and a request's details | RouterPlus as the provider, no host, the endpoint id as the model, and "contract rate" as the price | | The playground's receipt | served by RouterPlus | The records stay masked once an endpoint has been closed: what it served before it closed, and after it is opened again, still reads as above. The fields of a closed chat completion: | Level | Fields | |---|---| | Top level | `id` (`chatcmpl-` and the request id without dashes), `object`, `created` (when the gateway received your request), `model`, `choices`, `usage` | | A choice | `index`, `message` (not streamed) or `delta` (streamed), `finish_reason` (`stop`, `length`, `tool_calls`, `content_filter` or `function_call`; null until a stream ends), `logprobs` (only when you ask for them) | | `message` or `delta` | `role`, `content`, `tool_calls` (`id`, `type`, `function.name`, `function.arguments`, and `index` in a stream), `function_call`, `refusal`, `reasoning_content` | | `usage` | `prompt_tokens`, `completion_tokens`, `total_tokens`, `prompt_tokens_details.cached_tokens` and `completion_tokens_details.reasoning_tokens` (both always present, 0 when none), `cost` | Every other field is dropped, including any field a provider adds later, and a field without a value is left out (a message keeps `content`: null beside tool calls). A model's thinking arrives as `reasoning_content` on every route. A tool call's id is the gateway's own, `call_` and 24 hex characters, the same on every chunk of the call and as the `tool_use` id on `/v1/messages`. In a stream, the first delta of a choice carries the role, the first delta of a tool call its id, type and name, and the finish and the usage each come in a chunk of their own. ## Prices | Attempt | Price | |---|---| | On your dedicated capacity | your contract rate card, per million tokens | | On a fallback | the model's list price, as for any other request; or your rate card, when your contract sets one price on every route | A closed endpoint always bills your rate card on every route: a second price would tell you which route served. A fallback is tried only when its list price is known and at or below the endpoint's maximum price, if one is set, whichever price you then pay. In `/v1/generation`, an attempt billed at your rate card has `inference_price_source` `customer_rate` and a `price_snapshot_id` of the form `dedicated::r`, so every charge traces back to the rate card that produced it. `usage.cost` on the response is the full debit across all attempts, as on every request. ## Limits For the endpoint's traffic, its own limits replace the default organization, workspace, principal and key limits: | Limit | Counted for | |---|---| | requests per minute | the endpoint | | requests per second (optional): a count per second, or a paced rate | the endpoint | | tokens per minute | the endpoint | | concurrent requests | the endpoint | A limit your organization set itself on [`/console/limits`](https://app.routerplus.com/console/limits) still applies on top. You can lower what one key or workspace may use; you cannot raise the endpoint's limits. A refusal is the usual 429 `rate_limit` (see [Rate limits & spend caps](/docs/limits)), with `x-tm-limit-scope: endpoint`: | Case | `x-tm-limit-kind` | `x-tm-limit-id` | |---|---|---| | Over the endpoint's RPM, RPS, TPM or concurrency | `rpm`, `rps`, `tpm` or `concurrency` | `endpoint:` | | Above a paced endpoint's hard rate | `rps` | `endpoint:` | | A paced endpoint's wait would be longer than its bound | `rps_queue` | `endpoint:` | | The dedicated capacity is full and the endpoint does not spill over to the shared pool | the measure that is full | `endpoint:` | | A closed endpoint: the capacity behind every route it could take is full | the measure that is full | `endpoint:` | | Fallback traffic is over the endpoint's fallback cap | `rpm`, `tpm` or `concurrency` | `endpoint-fallback:` | `GET /v1/limits?model=/` returns your key's effective limits on the endpoint and a `dedicated` object: the endpoint, its revision, the model, the output default, the fallback caps and whether it spills over. If the endpoint's outage mode is `closed`, requests are refused while the shared counter store is unreachable. In the default `bounded_open` mode they are admitted under the endpoint's own limits, divided among the gateways that serve it, for up to 10 minutes; see [Limits and capacity](/docs/admission). ### The paced limit A limit that counts each calendar second refuses requests in many seconds when traffic arrives at random, even at an average below the limit. A paced limit queues a short burst instead. It has three settings: the paced rate, a hard rate above it, and the longest wait. - Requests leave admission no faster than the paced rate. A request that must wait waits in the gateway. `x-tm-queue-ms` on the response says how long, in milliseconds. - A request whose wait would be longer than the bound gets 429, `x-tm-limit-kind: rps_queue`. Its `retry-after` says when its wait would fit. - The hard rate refuses a request at once, `x-tm-limit-kind: rps`. It allows a burst as long as the bound (at least one second), so it never refuses a request the queue would hold: with a bound of a second or more, a full queue is what refuses. - The wait comes before the request's own deadline starts, so it does not use the endpoint's total time. - A waiting request holds no concurrency slot. Its place in the pace stays used even if you disconnect. For example, with a paced rate of 6 requests a second, a hard rate of 6.5 and a 2 s bound: | Average rate, arriving at random | Refused (429) | Typical wait | |---|---|---| | 4 a second | almost none | about 0.1 s | | 5 a second | about 0.3 % | about 0.3 s | | 6 a second | about 4 % | about 0.9 s | | Steadily above 6 a second | the excess, once the queue is full | up to 2 s | A queue cannot hold a steady excess: over time the endpoint passes its paced rate. ## Timeouts and defaults | Setting | On other requests | On an endpoint | |---|---|---| | Attempts per request | up to 8 | 1 to 8; 3 unless set | | Total time per request | 450 s | 1 to 600 s | | Wait for a streamed attempt's response headers before the next route | 20 s | 1 to 120 s | | Time one attempt may take on a request that is not streamed, before the next route | 450 s | 1 to 600 s, below the total | | `max_tokens` written into a request that sets none | 4,096 | set per endpoint | A shorter header wait, or attempt limit for requests that are not streamed, makes failover faster when the dedicated deployment stalls. A stalled request still takes that limit plus the fallback's own time. The attempt limit applies only while another route remains: the last route may run for the rest of the total time. ## Edges - `POST /v1/route` does not explain dedicated endpoints yet. - Health circuits are shared by every request to a deployment. Two failures in a row on the dedicated deployment send all its traffic to the fallbacks for 30 seconds. Per-endpoint circuit settings are not available yet. - We change an endpoint's configuration on request. Each change is a new revision, and it reaches every gateway within seconds. --- Source: https://app.routerplus.com/docs/models.md # Models & catalog The catalog is one flat namespace of model ids. Every **chat** model in it is callable through both wire surfaces — OpenAI-compatible `/v1/chat/completions` and Anthropic-compatible `/v1/messages` — with one key; the gateway translates. Image models answer on `POST /v1/images/generations` only, video models on `POST /v1/videos` only, and decision models on `POST /v1/decisions` only. Pricing is pass-through: the per-token rates in the catalog are what you pay, and every billed response carries `usage.cost` so you can recompute the bill from the wire. There are four ways to read the catalog: | Surface | URL | Auth | For | |---|---|---|---| | Model directory | `https://app.routerplus.com/` | none | Browsing, filters, 24-hour request counts | | Per-model pages | `https://app.routerplus.com/models/` | none | Providers with prices, latency, uptime and data policy; copy-paste snippets | | Machine-readable JSON | `https://app.routerplus.com/api/models.json` | none | Scripts, dashboards, price checks | | SDK-shaped list | `https://api.routerplus.com/v1/models` | API key | `client.models.list()` — see [GET /v1/models](/docs/api-models) | ## Model ids Ids are exact strings. There is no fuzzy matching and no silent substitution: a request for an id that is not in the catalog is a 404 that echoes the id you sent (see [GET /v1/models](/docs/api-models)). The catalog as of this writing — the live list is always [`/api/models.json`](https://app.routerplus.com/api/models.json). Models served by two providers list both, first-party first — see [One model, several providers](#one-model-several-providers): | Model id | Providers | Context | Input /1M | Output /1M | |---|---|---|---|---| | `claude-sonnet-5-5` | anthropic, openrouter | 1M | $2 | $10 | | `claude-opus-5-5` | anthropic, openrouter | 1M | $4 | $20 | | `claude-fable-5-1` | anthropic, openrouter | 1M | $10 | $50 | | `gpt-6.1-sol` | openrouter | 1.05M | $2 | $10 | | `gpt-6-sol` | openrouter | 1.05M | $2 | $10 | | `gpt-6-luna` | openrouter | 1.05M | $0.1 | $0.5 | | `gpt-6-astra` | openrouter | 1.05M | $10 | $50 | | `claude-sonnet-5` | anthropic, openrouter | 1M | $2 | $10 | | `claude-opus-5` | anthropic, openrouter | 1M | $5 | $25 | | `claude-fable-5` | anthropic, openrouter | 1M | $10 | $50 | | `claude-sonnet-4-5` | anthropic, openrouter | 200K | $3 | $15 | | `claude-haiku-4-5` | anthropic, openrouter | 200K | $1 | $5 | | `gpt-4o-mini` | openai, openrouter | 128K | $0.15 | $0.6 | | `gpt-4.1-mini` | openai, openrouter | 1.05M | $0.4 | $1.6 | | `gpt-4o` | openai, openrouter | 128K | $2.5 | $10 | | `tencent/hy4-preview` | openrouter | 1.05M | $0.834 | $2.501 | | `z-ai/glm-5.3-flash` | openrouter | 1.31M | $0.075 | $0.25 | | `deepseek/deepseek-v4.1-flash` | openrouter | 1.05M | $0.3 | $1.2 | | `deepseek/deepseek-v4-flash-0731` | openrouter | 1.31M | $0.04998 | $0.09996 | | `deepseek/deepseek-v4-flash` | openrouter | 1.05M | $0.08078 | $0.16156 | | `tencent/hy3` | openrouter | 262K | $0.132 | $0.528 | | `z-ai/glm-5.3` | openrouter | 1.31M | $1.4 | $4.4 | | `xiaomi/mimo-v2.5` | openrouter | 1.05M | $0.14 | $0.28 | | `z-ai/glm-5.2` | openrouter | 1.05M | $0.966 | $3.036 | | `google/gemini-3.7-flash` | openrouter | 1.05M | $0.75 | $3.75 | | `moonshotai/kimi-k3` | openrouter | 1.05M | $3 | $15 | | `minimax/minimax-m3` | openrouter | 1.05M | $0.3 | $1.2 | | `deepseek/deepseek-v4-pro` | openrouter | 1.05M | $0.687648 | $1.375296 | | `upstage/solar-pro4` | openrouter | 524K | $0.03 | $0.12 | | `deepseek/deepseek-v4-pro-0813` | openrouter | 1.05M | $1.12068 | $3.36204 | | `google/gemini-3.8-flash` | openrouter | 1.05M | $0.75 | $3.75 | A model or mix you deploy from [Optimize](/docs/model-search) gets an id of its own, `tm/-v`. It is callable on `POST /v1/chat/completions` with a key of the organization that owns it, and it is not listed in the catalog. ### Image models Image models bill in tokens like everything else, but they answer on [POST /v1/images/generations](/docs/api-images) and nowhere else. The per-image ceiling is the most output tokens one image can bill; the reservation holds `n` times it. | Model id | Sizes | Qualities | Per-image ceiling | Text in /1M | Image out /1M | |---|---|---|---|---|---| | `gpt-image-1` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high` | 6,240 | $5 | $40 | | `gpt-image-1-mini` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high` | 6,500 | $2 | $8 | | `gpt-image-1.5` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high` | 6,500 | $5 | $32 | | `gpt-image-2` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high` | 8,232 | $5 | $30 | | `gpt-image-2.5-flare` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high · xhigh · max` | 8,232 | $5 | $30 | | `gpt-image-2.5-sunburst` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high · xhigh · max` | 8,232 | $5 | $30 | `size` and `quality` also accept `auto`, which lets the provider choose. Every model's full descriptor map is in `/api/models.json` under `supported_parameters`. The Nano Banana, Seedream and Grok Imagine models are served through OpenRouter and billed at the cost it reports ([how](/docs/pricing#models-billed-at-the-providers-reported-cost)). They take a shape as `aspect_ratio` (`16:9`, `9:16`, …) instead of a `size`, most take a `resolution` (`1K`, `2K`, `4K`), and Seedream takes a `seed`. ### Video models Video models answer on [POST /v1/videos](/docs/api-videos) and nowhere else. A video is a job: it is created at once and made in the next seconds to minutes. All three are served through OpenRouter today and billed at the cost it reports. | Model id | Made by | Length | Resolutions | Listed price | Per second ≤ | |---|---|---|---|---|---| | `seedance-2.5` | ByteDance | 4–30 s | 480p · 720p | $10.70 / M video tokens | $0.30 | | `hailuo-3-max` | MiniMax | 5–15 s | 480p · 768p | $0.05–$0.08 per second | $0.08 | | `wan-3.0` | Alibaba | 5–30 s | 480p · 720p · 1080p | $0.05–$0.20 per second | $0.20 | ### Decision models Decision models answer on [POST /v1/decisions](/docs/api-decisions) and nowhere else. They write no text: they return one typed answer per question about a state. Each model takes its own question format. | Model id | Made by | Providers | Question kinds | Input /1M | |---|---|---|---|---| | `typesafe/jev-1.13` | TypeSafe | `typesafe`, then `openrouter-decisions` | choice, score, yes/no (System One) | $0.042 | | `inception/mercury-decide` | Inception | `openrouter-decisions-free` | choice, score, yes/no (System One, the same as Jev) | $0 | | `bespokelabs/nimble-v3` | Bespoke Labs | `bespokelabs` | choice, score, yes/no (System One, the same as Jev) | $0.04 | | `cloudflare/clef` | Cloudflare | `workers-ai` | choice, score, yes/no (System One, the same as Jev) | $0.24 | | `cloudflare/clef-flash` | Cloudflare | `workers-ai` | choice, score, yes/no (System One, the same as Jev) | $0.09 | | `routerplus/decider-2b` | RouterPlus | `routerplus` | choice, score, yes/no (System One, the same as Jev) | $0 | | `routerplus/kev-4b` | RouterPlus | `routerplus` | choice, score, yes/no (System One, the same as Jev) | $0 | | `perplexity/pplx-decider-v1-27b` | Perplexity | `perplexity-decisions` | choice, score, yes/no (System One, the same as Jev) | $0.04 | | `levanto/sage-1.2` | Levanto | `levanto` | choice, score, yes/no (System One, the same as Jev) | $0.05 | | `fastino/gliner-2.5-decide` | Fastino | `fastino` | classifications (single- and multi-label), entities, structures, relations | $0.03 | Output is free on every decision model. Sage, like Perplexity Decider, bills the state once per question. Mercury Decide is $0 in and out while OpenRouter serves only its free variant; its pool takes 20 requests a minute (15 per organization) and 5 in flight at once (3 per organization), and OpenRouter caps free requests per day. Bespoke Nimble v3 is $0.04 per million input tokens, Bespoke's own price with no markup. It takes 1 to 64 questions per call, and its pool takes 7 requests at once (5 per organization): Bespoke allows our account 8, and one is kept for our test environment. Clef and Clef-flash are Cloudflare's own decision models, served by Cloudflare Workers AI at Cloudflare's own price with no markup: $0.24 and $0.09 per million input tokens. They take 1 to 64 questions per call, each question id 1 to 100 letters, digits, `_`, `.` or `-`, and a context of 65,536 tokens. They share one pool: 200 requests a minute (150 per organization) and at most 8 at a time (6 per organization). Decider 2B and Kev 4B are RouterPlus's own models, on our own GPUs in the United States, and free: $0 in and out. They share one pool. Their contexts are 25,600 and 8,192 tokens. RouterPlus answers an identical request from its cache for 600 s. Perplexity Decider v1 27B is Perplexity's own decision model, served by Perplexity's API at Perplexity's own price with no markup: $0.04 per million input tokens. Perplexity bills the state once per question, so a call costs about one state for each question it asks. It takes 1 to 128 questions per call, and 262,144 tokens for each question's prompt. Its pool takes 540 requests a minute (405 per organization) and 8 at once (6 per organization). A discount may apply to decision models; see [Pricing](/docs/pricing). Most models also carry cache pricing (`cached_prompt`, and `cache_write` where the provider bills it) — the full per-SKU table is on each model's page and in `/api/models.json`. A model with no separate cache or reasoning price bills those tokens **at the base rate** — a missing SKU never means free. ### One model, several providers A model id can be served by more than one provider — a first-party lab and an aggregator, say. The catalog shows one card; the model page lists every provider with its own uptime and TTFB. Routing tries providers in priority order and fails over **before the first byte** only — never mid-answer. Prices are identical across providers of one model (enforced when the catalog is built), so which provider served you never changes your bill. The `x-tm-provider` header says who did; when that provider knows the model under its own id (an aggregator's `anthropic/claude-sonnet-4.5` for our `claude-sonnet-4-5`), `x-tm-upstream-model` carries the id that was sent and the response's `model` field echoes the provider's id truthfully. When an aggregator names the provider it used, `x-tm-served-by` carries that name (for example `Amazon Bedrock`). ### Date-pinned ids The catalog lists **one clean id per model** (`claude-haiku-4-5`), but vendors also ship dated snapshots (`claude-haiku-4-5-20251001`, `gpt-4o-2024-08-06`) and `-latest` spellings, and tools like Claude Code send them. Any such pin of a catalog model works without being listed: the gateway resolves it to its family **for routing and pricing only**, and forwards your original id to the provider verbatim — a snapshot pin is never silently retargeted to a newer version. The response's `model` field shows the exact snapshot that served you. If a vendor retires a pinned snapshot, you get the vendor's own error, not a quiet upgrade. Unknown ids still 404 naming themselves, dated or not. A key bound to your own provider connection matches ids literally, so list the exact ids you send there. ### Id grammar Ids match `[a-zA-Z0-9][a-zA-Z0-9._/-]{0,127}` — letters, digits, `.` `_` `/` `-`, max 128 chars. `:` is deliberately excluded: the suffix namespace (`:something`) is reserved for future marketplace variants, so a provider can never squat a routing suffix. ## Providers are manifests A provider is a JSON manifest in the repo — endpoint, wire dialect, settlement terms, data policy, and a model list with exact prices: ```json { "provider": { "id": "anthropic", "name": "Anthropic (direct API)", "prompt_logging": "retained" }, "endpoint": { "base_url": "https://api.anthropic.com/v1", "dialect": "anthropic" }, "models": [{ "id": "claude-haiku-4-5", "context_length": 200000, "max_output_tokens": 64000, "pricing": [ { "type": "prompt", "unit": "token", "cost_usd_per_million": "1" }, { "type": "completion", "unit": "token", "cost_usd_per_million": "5" } ] }] } ``` Manifests compile into the routing catalog through a gated pipeline — validate, emit a staging catalog, run the live conformance suite, then promote. Adding or repricing a model is a **data change with a gate**, never a gateway code deploy. Three lifecycle rules do real work: - **`is_ready: false`** stages a model: validated and testable, but never routed and never listed. - **`deprecation_date`** delists automatically: past that date the model drops out of the compiled catalog and the public pages. - **Shared ids must agree.** If two providers declare the same model id with different prices, or one as a chat model and the other as an image model, the catalog build fails. Pass-through pricing stays coherent per id. > [!NOTE] > Prices sync as effective-dated append-only rows. A price change never > rewrites history — old requests stay priced as dispatched. ## Conformance before listing No model is listed on a provider's say-so. The conformance suite runs live requests against the provider's endpoint, and passing it is the gate between staging and live. It is deliberately implemented independently of the gateway's own wire parsers, so a shared bug can't hide real breakage. | Check | Proves | |---|---| | C1 | A non-stream completion returns real usage tokens (usage is the billable record) | | C2 | Stream frames are parseable SSE | | C3 | Streams terminate properly (`[DONE]` / `message_stop`) | | C4 | Usage tokens are reported in-stream | | C5 | An unknown model yields a parseable 4xx — not a silent fallback | | C6 | Responses carry actual assistant content, not just billable counters | | C7 | The response echoes the requested model id — no silent substitution | | I1 | An image generation at the top declared quality returns base64 data and usage | | I2 | The declared per-image ceiling covers the billed output | | I3 | An undeclared size is refused | | I4 | The ceiling holds at every declared size (release-day probe, opt-in: it buys a real render per size) | | V1 | A video job at the cheapest tier finishes, reports its cost and downloads as an MP4 | | V2 | The declared per-second ceiling covers the reported cost | | V3 | An undeclared shape is refused | | V4 | The ceiling holds at the top resolution (opt-in: it buys one more job per model) | Image listings run I1–I3 and C5; video listings run V1–V3 and C5. C2–C4, C6 and C7 do not apply to them — the media routes do not stream, carry no assistant text, and an images response has no `model` field to echo. C6 and C7 exist because the failure modes that matter are billing for an empty answer and quietly serving a cheaper model. C7 accepts exactly one alias form: a dated snapshot of the *same* id (`gpt-4o` → `gpt-4o-2024-08-06`), or an alias target the manifest declares up front in `resolves_to`. `gpt-4o` answered by `gpt-4o-mini` fails. ## The directory [`https://app.routerplus.com/`](https://app.routerplus.com/) lists every live model in three sections — chat, image and video — because they are called on different routes. A row shows the model's context length and max output (for a chat model), or its sizes, qualities, lengths and resolutions (for an image or video model), plus its 24-hour request count. The list can be searched (`/?q=haiku` matches model, lab and provider names), filtered by kind, context length, supported parameters, zero data retention, input price, lab and provider, and sorted by use, name, price or context. Prices, latency and uptime depend on the provider, so they live on each model's page at `/models/`: one row per provider that runs the model, per-SKU pricing, the provider's prompt-retention policy (marketplace logs are content-free either way), and ready-to-paste snippets. ### The uptime figure is allowed to say nothing Uptime and latency come from the trailing 24 hours of real routed traffic. When a provider we call directly has fewer than 100 counted attempts in that window, the model page shows **`n<100`** instead of a percentage. For a provider reached through an aggregator, the figure is the one the aggregator publishes. > [!NOTE] > A model with 3 requests is not "100% up", and we won't render it that way. > No number until the sample is real — that rule is load-bearing, not a > placeholder. The hourly bars on a model page are green at ≥ 99%, amber at ≥ 95%, and red below that; an hour with no sample stays grey. ## Machine-readable: /api/models.json Everything above, as JSON, no auth: ```bash curl -s https://app.routerplus.com/api/models.json ``` ```json { "models": [ { "id": "claude-haiku-4-5", "display_name": "Claude Haiku 4.5", "output_modalities": ["text"], "providers": [ { "id": "anthropic", "name": "Anthropic (direct API)", "dialect": "anthropic", "prompt_logging": "retained" }, { "id": "openrouter", "name": "OpenRouter", "dialect": "openai", "prompt_logging": "retained", "upstream_id": "anthropic/claude-haiku-4.5" } ], "served_by": [ { "id": "anthropic", "name": "Anthropic", "zdr": false, "routes": [ { "via": "anthropic", "direct": true }, { "via": "openrouter", "direct": false } ] }, { "id": "amazon", "name": "Amazon", "zdr": true, "routes": [ { "via": "openrouter", "direct": false } ] }, { "id": "azure", "name": "Azure", "zdr": false, "routes": [ { "via": "openrouter", "direct": false } ] }, { "id": "google-vertex", "name": "Google Vertex", "zdr": true, "routes": [ { "via": "openrouter", "direct": false } ] } ], "context_length": 200000, "max_output_tokens": 64000, "pricing_usd_per_million": { "prompt": "1", "cached_prompt": "0.10", "cache_write": "1.25", "completion": "5" }, "lab": { "id": "anthropic", "name": "Anthropic" }, "weights": "proprietary" } ], "gateway": "https://api.routerplus.com/v1", "docs": "https://app.routerplus.com/docs" } ``` | Field | Meaning | |---|---| | `id` | The exact string to put in your request's `model` field | | `output_modalities` | `["text"]`, `["image"]`, `["video"]` or `["decisions"]` — which route serves this model. Filter on it; never guess from the id | | `providers` | Every deployment serving this id, primary first, each with `id`, `name`, `dialect`, `prompt_logging`, and `upstream_id` when the provider knows the model under its own name | | `dialect` (per provider) | The provider's **native** wire format — informational; every chat model answers on both surfaces | | `prompt_logging` (per provider) | Provider's prompt retention: `none` or `retained` | | `served_by` | The providers that run the model, one entry per provider however it is reached. `routes[].via` is the `providers` entry the request goes to; `direct: false` means an aggregator picks this provider per request. `quantizations` appears when the aggregator publishes it (for example `["fp8"]`). `zdr` is `true` when the provider offers Zero Data Retention for this model | | `context_length` / `max_output_tokens` | Token limits; `max_output_tokens` may be `null`, and `context_length` is `null` for image and video models | | `supported_parameters` | Image and video models only: the descriptor map the route enforces (`size`, `quality`, `n`, `seconds`, and the rest) | | `pricing_usd_per_million` | Decimal-string USD per 1M tokens, keyed by SKU type | | `billing` | `"reported_cost"` on the models billed at the provider's reported cost; absent on token-billed models | | `price_card` | On those models, the provider's own listed prices, for reading — see [Pricing & billing](/docs/pricing#models-billed-at-the-providers-reported-cost) | | `lab` | Who made the model, as distinct from who serves it — `null` when unknown | | `weights` | `"open"` when the lab publishes the model's weights, `"proprietary"` when it never has, `null` when we are not certain | For an image model `max_output_tokens` is the **per-image ceiling**, not a cap on a completion. For a video model it is the ceiling per second of video, in cost units (one micro-dollar each). Prices are decimal **strings** so nothing is lost to float rounding — parse them with a decimal type if you're doing money math. For the authenticated, SDK-shaped list (what `client.models.list()` calls), see [GET /v1/models](/docs/api-models). --- Source: https://app.routerplus.com/docs/pricing.md # Pricing & billing Pricing is pass-through: the per-token price you pay is the provider's price, published per model in USD per million tokens. There is no per-token markup and no hidden fee line. Every billed response carries `usage.cost` — the exact amount your balance was debited, computed with the same integer math the ledger settles with — so you can recompute your bill from the wire at any time. One provider is our own: RouterPlus, which serves our own decision models on our own GPUs. Its models are free: $0 in and out (see [Decision models](#decision-models)). ## Where prices live The machine-readable catalog is public, no key required: ```bash curl -s https://app.routerplus.com/api/models.json \ | jq '.models[] | select(.id == "claude-sonnet-4-5").pricing_usd_per_million' ``` ```json { "prompt": "3", "cached_prompt": "0.30", "cache_write": "3.75", "completion": "15" } ``` Prices are **decimal strings** (up to six fractional digits), USD per one million tokens. They convert exactly to the ledger's internal unit — no float rounding is involved anywhere in billing. | SKU | What it bills | Example (`claude-sonnet-4-5`) | |---|---|---| | `prompt` | non-cached input tokens | $3 / M | | `cached_prompt` | prompt-cache **reads** | $0.30 / M | | `cache_write` | prompt-cache **writes** | $3.75 / M | | `completion` | output tokens (including reasoning, unless a distinct rate is declared) | $15 / M | | `internal_reasoning` | reasoning subset of output, when a provider declares a distinct rate | — | > [!NOTE] > A model that does not declare a cache or reasoning SKU bills those tokens at > its base rate: `cached_prompt` and `cache_write` default to the `prompt` > price, `internal_reasoning` defaults to the `completion` price. A missing > price never means free — that would be an accidental $0 SKU, and the ledger > refuses to guess. ## Image models Image models bill on the **same two SKUs**: `prompt` for the text you send in, `completion` for the image that comes out, both in USD per million tokens. OpenAI reports image output in tokens, and the count is fixed by the size and quality you ask for. The per-image ceiling — the most output tokens one image can bill — is `max_output_tokens` on the model page and in `/api/models.json`. | Model id | Text in /1M | Image out /1M | Per-image ceiling | |---|---|---|---| | `gpt-image-1` | $5 | $40 | 6,240 tokens | | `gpt-image-1-mini` | $2 | $8 | 6,500 tokens | | `gpt-image-1.5` | $5 | $32 | 6,500 tokens | | `gpt-image-2` | $5 | $30 | 8,232 tokens | | `gpt-image-2.5-flare` | $5 | $30 | 8,232 tokens | | `gpt-image-2.5-sunburst` | $5 | $30 | 8,232 tokens | Every output token bills at the image rate, including the text output tokens `gpt-image-1.5` folds into its `output_tokens`. The full route contract is [POST /v1/images/generations](/docs/api-images). ## Models billed at the provider's reported cost Some image and video models are served through OpenRouter today: Google's Nano Banana family, ByteDance's Seedream and Seedance, xAI's Grok Imagine, MiniMax's H3 Max and Alibaba's Wan 3.0. Their providers price them per image, per second of video or per video token — not in text tokens. For these models **the charge is exactly the cost the provider reports** for the request, passed through with no margin. - The models page and `/api/models.json` show each model's own listed prices (`price_card`). - The reservation is a ceiling: per image for an image model, per second for a video model. The models page shows it as **Per image ≤** or **Per second ≤**. - When the provider reports no cost, the charge is the ceiling, marked `estimated`. - A reported cost above the ceiling is still charged in full: the ceiling limits what is held, not what is billed. In the ledger these rows count *cost units*: one unit is one micro-dollar, priced at $0 per million in and $1 per million out. So `output_tokens` on such a row is the charge in micro-dollars, and the per-million prices in `/api/models.json` read `"0"` and `"1"` for these models. The response you read still carries the provider's own token counts, and `usage.cost` in USD. | Model id | Kind | Listed price | Ceiling | |---|---|---|---| | `gemini-2.5-flash-image` | image | $30 / M image tokens | $0.05 per image | | `gemini-3.1-flash-image` | image | $60 / M image tokens | $0.20 per image | | `gemini-3-pro-image` | image | $120 / M image tokens | $0.30 per image | | `seedream-5.0-pro` | image | $0.045 per image; $0.09 at 2K | $0.10 per image | | `seedream-5.0-lite` | image | $0.035 per image | $0.04 per image | | `grok-imagine-image-2.0` | image | $0.04–$0.08 per image, by quality and resolution | $0.09 per image | | `seedance-2.5` | video | $10.70 / M video tokens | $0.30 per second | | `hailuo-3-max` | video | $0.05 per second at 480p, $0.08 at 768p | $0.08 per second | | `wan-3.0` | video | $0.05, $0.10, $0.20 per second at 480p, 720p, 1080p | $0.20 per second | The route contracts are [POST /v1/images/generations](/docs/api-images) and [POST /v1/videos](/docs/api-videos). ## Decision models Decision models answer typed questions on [POST /v1/decisions](/docs/api-decisions). They are priced like everything else: USD per million tokens. | Model id | Input /1M | Output /1M | What the token counts are | |---|---|---|---| | `typesafe/jev-1.13` | $0.042 | $0 | The provider's input count | | `inception/mercury-decide` | $0 | $0 | The provider's input count | | `bespokelabs/nimble-v3` | $0.04 | $0 | The provider's input count | | `cloudflare/clef` | $0.24 | $0 | The provider's input count | | `cloudflare/clef-flash` | $0.09 | $0 | The provider's input count | | `routerplus/decider-2b` | $0 | $0 | The provider's input count; an answer from its cache bills $0 | | `routerplus/kev-4b` | $0 | $0 | The provider's input count; an answer from its cache bills $0 | | `perplexity/pplx-decider-v1-27b` | $0.04 | $0 | The provider's input count, which holds the state once per question | | `levanto/sage-1.2` | $0.05 | $0 | The provider's input count, which holds the state once per question | | `fastino/gliner-2.5-decide` | $0.03 | $0 | Fastino's `prompt_tokens` | Mercury Decide is $0 in and $0 out because OpenRouter, its only provider, serves only the model's free variant today, and pricing is pass-through. `usage.cost` is `0` on every call. A paid variant, or Inception serving the model directly, would be a new listed price, shown on the model's page and in `/api/models.json` like any other. A free call is still metered, counted against your limits and its pool, and recorded in `GET /v1/generation`, and it still needs an organization with credit (see the discount note below). Bespoke Nimble v3 is billed on Bespoke's own input count at Bespoke's own price, with no markup: $0.04 per million input tokens, and output is free. Bespoke counts the state and the questions once per call, not once per question. A three-question call of about 340 input tokens costs $0.000013. Clef and Clef-flash are billed on Cloudflare's own input count at Cloudflare's own price, with no markup: $0.24 and $0.09 per million input tokens, and output is free (Cloudflare reports `output_tokens: 0`). Cloudflare bills Workers AI in neurons, at $0.011 per 1,000 neurons. Its price page lists Clef at 21,818 neurons ($0.24) and Clef-flash at 8,182 neurons ($0.09) per million input tokens. A call of 400 input tokens costs $0.000096 on Clef and $0.000036 on Clef-flash. Decider 2B and Kev 4B are RouterPlus's own models, on our own GPUs, not an outside provider's. So their price is ours to set, not a provider's price passed through, and it is $0: input, cached input and output are all free, and `usage.cost` is `0` on every call. RouterPlus still reports its input count, which is at most the model's context: a long input that Decider 2B (25,600 tokens) cuts counts only the tokens the model read. A free call is still metered, counted against your limits and its pool, and recorded in `GET /v1/generation`, and it still needs an organization with credit. RouterPlus answers an identical request from its cache for 600 s. Perplexity Decider v1 27B is billed on Perplexity's own input count at Perplexity's own price, with no markup: $0.04 per million input tokens, and output is free. Unlike Bespoke, Perplexity runs one prompt for each question, each with the whole state, and counts the state once per question. So a call costs about one state for each question it asks: five questions about a state of 1,843 tokens bill 9,215 input tokens, $0.000368. One call with many questions still saves round trips. GLiNER-2.5-Decide is billed on Fastino's own input count: $0.03 per million input tokens, and output is free. `usage.input_tokens` is Fastino's `prompt_tokens`, and `usage.output_tokens` its `completion_tokens`. A call of 1,000 input tokens costs $0.00003. Sage is billed on Levanto's own input count at Levanto's own price, with no markup: $0.05 per million input tokens, and output is free. Like Perplexity Decider, it reads and bills the state once per question, so a call costs about one state for each question it asks. ### The decisions discount A discount may apply to decision models: a percent taken off the charge of every decision model, or of some of them, each at its own percent, from every key, the playground's included. It does not change the list price, and it can change or end. A discounted model's page and its row in the directory show the percent beside the struck list price, and its row in `/api/models.json` carries `discount_percent`. Every discounted response says what applied: | Where | What it shows | |---|---| | `usage.cost` | What you paid, after the discount | | `usage.cost_before_discount` | The charge at the list price | | `usage.discount_percent` and the `x-tm-discount-percent` header | The percent taken off | | `cost_usd` in [GET /v1/generation](/docs/api-usage) | What you paid, after the discount | | `inference_cost_usd` in GET /v1/generation | The cost at the list price | The discounted charge is the list charge × (100 − percent) / 100, rounded down to the micro-dollar. When no discount applies, the three discount fields are absent and `usage.cost` is the list charge. A discount is a lower charge, not a markup: pricing stays pass-through. A discount is for organizations with credit: with a balance of $0 a discounted request is refused with `429 insufficient_quota`, as any request is, even at 100 %. When a discount covers a model whose list price is already $0 (Mercury Decide, Decider 2B and Kev 4B today), the discount has nothing to take off: `usage.cost` and `usage.cost_before_discount` are both `0`. The gateway still reports `usage.discount_percent` and the header on it, as on every model the discount covers, and `/api/models.json` still carries `discount_percent` on its row, so one reader works for all ten models. The model's page, the directory and the Playground show no "N % off" beside a $0 price: the model is free, not discounted. ## The math: integer micro-USD All money math is integer micro-USD (millionths of a dollar). A published price like `"2.50"` becomes exactly `2500000` micro-USD per million tokens. Two rounding rules, both fixed: - **Reserve rounds up** (ceiling) — the pre-flight hold is conservative. - **Settlement rounds down** (floor) — the actual charge, in your favor. The settlement formula, verbatim from the ledger: ``` non_cached_input = max(0, input_tokens − cached_tokens − cache_write_tokens) reasoning = clamp(reasoning_tokens, 0, output_tokens) plain_output = output_tokens − reasoning cost_micro = floor(( non_cached_input × prompt_per_M + cached_tokens × cached_prompt_per_M + cache_write_tokens × cache_write_per_M + plain_output × completion_per_M + reasoning × reasoning_per_M ) / 1,000,000) ``` where each `*_per_M` is the integer micro-USD price. `usage.cost` on the wire is this number divided by 10⁶. ## Reserve, then settle Every billed request runs the same lifecycle: 1. **Reserve.** Before any provider is contacted, the gateway holds a conservative worst case. On the chat surfaces the input is estimated at one token per byte of the request body, plus 65,536 tokens for each image or document part (an image sent inline counts the 65,536 only, not its base64 bytes), and the output at your `max_tokens` (or `max_completion_tokens`; 4096 when omitted, and never more than 32,768 — see [POST /v1/chat/completions](/docs/api-chat-completions)). Each side is priced at the model's most expensive rate for that side — cache writes and reasoning when they cost more than plain input or output — and the sum is rounded **up** to micro-USD. On `/v1/images/generations` the output is reserved at `n` × the model's per-image ceiling (not `max_tokens`), and the input at the prompt bytes ÷ 4; on `/v1/videos` the hold is the length in seconds × the model's per-second ceiling. 2. **Dispatch.** The request goes to a provider. Prices are pinned at this moment (see [frozen snapshots](#price-changes-frozen-snapshots)). 3. **Settle.** When the attempt finishes, the provider-reported token counts are priced with the formula above, rounded **down**. That is what you pay. 4. **Release.** The hold is released in full. You are never charged the reserved maximum — with one honest exception: if the gateway crashes mid-request and cannot observe the outcome, the attempt settles at the reservation, because the call may have been billed upstream. That outcome is visible as `unknown_after_crash` in [GET /v1/usage & /v1/generation](/docs/api-usage). A reservation can be refused before dispatch: - **Balance too low** — your balance must cover every reservation your org has in flight, including this one. Refusal is a `429` with `error_type: insufficient_quota` (the OpenAI-native shape SDKs already handle; `billing_error` native type on the Anthropic surface). - **Monthly spend cap hit** — hard caps (org-wide and per-key, UTC calendar month, strictest wins) return the same `429 insufficient_quota` with the exact reset time in the message and the `x-tm-cap-reset` header. - **A rate limit hit** — the key's RPM, a shared limit of the organization, or the provider's pool: `429` with `error_type: rate_limit` and `retry-after`. Remediation for each is in [Errors & remediation](/docs/errors). ### All-or-nothing on images An image generation either arrives whole or costs nothing. A failed attempt, a response over the 32 MiB cap, and a 200 that carries no decodable image all settle at zero — there is no partial image and no partial bill. The one nuance: if you hang up mid-render the render still finishes upstream, so that attempt settles the **provider-reported** usage, never the reservation and never a pretend $0. ## The cost on the wire: `usage.cost` Every billed response carries `usage.cost` (USD, a JSON number): - **Non-streaming** — inline in the response body's `usage`, both surfaces. - **Streaming, OpenAI surface** — in the final usage chunk, before `[DONE]`. - **Streaming, Anthropic surface** — in the `message_delta` usage at stream end. ```bash curl -s https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"claude-sonnet-4-5","max_tokens":350,"messages":[{"role":"user","content":"hello"}]}' \ | jq .usage ``` ```json { "prompt_tokens": 1200, "completion_tokens": 350, "total_tokens": 1550, "prompt_tokens_details": { "cached_tokens": 800, "cache_write_tokens": 0 }, "cost": 0.00669 } ``` ```python import os from openai import OpenAI client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) r = client.chat.completions.create( model="claude-sonnet-4-5", max_tokens=350, messages=[{"role": "user", "content": "hello"}], ) print(r.usage.model_dump()["cost"]) # the exact ledger debit, in USD ``` ```python import os import anthropic client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"]) msg = client.messages.create( model="claude-sonnet-4-5", max_tokens=350, messages=[{"role": "user", "content": "hello"}], ) print(msg.usage.model_dump()["cost"]) ``` > [!WARNING] > `usage.cost` is the **full request debit**, not the serving attempt's slice. > If a provider failed mid-prompt and the gateway failed over before your > answer started, any tokens the failed attempt consumed are included in the > number you see. The per-attempt breakdown is always available at > [GET /v1/generation](/docs/api-usage) — nothing is hidden in an average. One documented limitation: an Anthropic-surface **stream that dies after your answer started** has no legal wire slot for usage outside `message_delta`, so its terminal error event carries no cost. Recompute those from [/v1/generation](/docs/api-usage). (The OpenAI surface emits any known usage+cost before its terminal error, so its wire stays complete.) Usage reporting is always on: streams and non-streams both carry it, and the gateway requests usage upstream regardless of what you send, because unreported usage would be an unrecomputable bill. The one `stream_options` value that changes anything is an explicit `include_usage: false`: you are billed the same, but no usage chunk is written to you. ## Cache reads and writes Cache tokens are billed at their own rates on **both** surfaces, and both surfaces report the split: | Concept | OpenAI surface | Anthropic surface | Billed at | |---|---|---|---| | Non-cached input | `prompt_tokens` minus the two details below | `input_tokens` | `prompt` | | Cache reads | `prompt_tokens_details.cached_tokens` | `cache_read_input_tokens` | `cached_prompt` | | Cache writes | `prompt_tokens_details.cache_write_tokens` | `cache_creation_input_tokens` | `cache_write` | | Output | `completion_tokens` | `output_tokens` | `completion` | On the OpenAI surface `prompt_tokens` is cache-inclusive (reads and writes are inside it), and billed responses always carry `prompt_tokens_details` with both fields — zero-filled when the upstream reports none — so stream and non-stream usage share one recomputable shape. ## Reasoning tokens Reasoning tokens are a **subset of output**, never billed on top of it. The settle math splits them out and prices them at the model's reasoning rate; today every cataloged model prices reasoning equal to `completion`, so the split is billing-neutral until a provider declares a distinct rate. The OpenAI surface reports them in `completion_tokens_details.reasoning_tokens`; the Anthropic wire does not report them separately. ## Price changes: frozen snapshots The price you pay is the price at **dispatch time**. Price rows are append-only — a change is a new effective-dated row, never an edit — and the row in effect when your request is reserved is pinned into the reservation and written onto every attempt's ledger row. A price change published while your request is in flight cannot touch it, and historical attempts always settle (and audit) against the prices they were dispatched under. ## Paid credits and purchase matching Create an account at [Sign up](https://app.routerplus.com/signup), complete Clerk authentication and email verification, and save your first API key. Signup and card verification alone do not grant credit. Existing customers keep their balance and API keys. Standard credit purchases start at **$10** in [Billing](https://app.routerplus.com/console/billing). The payment adds the amount you buy as paid credit. Your balance is the sum of credit ledger entries minus settled spend. We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match. With the full match available: $10 purchase → $20 credits. $100 purchase → $200 credits. A partially used cap adds only the matching credit remaining. The cap applies across eligible purchases, not afresh to each checkout. Billing shows the offer and remaining matching credit for your account. Matching credit is promotional and appears separately in the ledger. Refunds and chargebacks reverse the corresponding match; a partial refund reverses its proportional matching credit. Adding a card without a purchase does not earn a bonus. Existing trial and card-verification grants stay in the ledger, but those signup offers are no longer available. Before an organization buys credit, its customer keys share the trial rate of 20 requests/minute. A purchase lifts that cap; each key then has its own ceiling (300 for a console key). Historical trial credit can also be limited when unusual usage is detected; contact support if that happens. ## Worked example A `claude-sonnet-4-5` request reports 1,200 prompt tokens (800 of them cache reads, no cache writes) and 350 completion tokens: ``` non-cached input: 400 × 3,000,000 = 1,200,000,000 cache reads: 800 × 300,000 = 240,000,000 output: 350 × 15,000,000 = 5,250,000,000 ───────────── sum / 1,000,000 (floor) = 6,690 micro-USD ``` `usage.cost` on the wire: `0.00669`. The same number appears as the attempt's `cost_usd` in [GET /v1/usage & /v1/generation](/docs/api-usage), because they are the same computation on the same integers. ## Worked example: an image A `gpt-image-1` request for two medium 1024×1024 images (`n: 2`) with a 12-token prompt estimate. Text in is $5 / M, image out is $40 / M, and the per-image ceiling is 6,240 tokens. The reservation, before any provider is contacted: ``` prompt estimate: 12 × 5,000,000 = 60,000,000 output ceiling: 2 × 6,240 × 40,000,000 = 499,200,000,000 ─────────────── ceil(sum / 1,000,000) = 499,260 micro-USD ``` The provider reports 12 input tokens and 2,112 output tokens (1,056 per image), so settlement is: ``` input: 12 × 5,000,000 = 60,000,000 output: 2,112 × 40,000,000 = 84,480,000,000 ────────────── floor(sum / 1,000,000) = 84,540 micro-USD ``` `usage.cost` on the wire: `0.08454`. The hold is released in full; you pay the settled number, not the reservation. --- Source: https://app.routerplus.com/docs/streaming.md # Streaming Set `"stream": true` and both surfaces stream Server-Sent Events (SSE): the OpenAI-compatible surface at `POST https://api.routerplus.com/v1/chat/completions` and the Anthropic-compatible surface at `POST https://api.routerplus.com/v1/messages`. The images route (`POST https://api.routerplus.com/v1/images/generations`) does not stream: `stream: true` and `partial_images` are typed 400s there. A video is a job you poll, not a stream — see [POST /v1/videos](/docs/api-videos). **Same-dialect routes relay.** When the provider serving your model speaks your surface's dialect (a Claude model on `/v1/messages`, a GPT model on `/v1/chat/completions`), the provider's stream is relayed record-for-record with payloads untouched — provider-specific features (server tool use, web search results, citations, annotations, `pause_turn`, the matched `stop_sequence`) reach you exactly as if you called the provider directly. The gateway only observes (for billing and failover), injects `usage.cost` into the provider's own usage record, and guarantees a well-formed terminator. One exception: an explicit `stream_options: {"include_usage": false}` turns the relay off, because the gateway must then remove the usage chunk it asked the provider for. **Cross-dialect routes re-encode.** When the provider speaks the other dialect, the gateway translates its stream into your surface's exact wire framing, and the shapes below hold. Before the first byte there is a 20-second headers deadline — a provider that can't answer in time is failed over invisibly (see [Routing & failover](/docs/routing)). After that a stream is not cut for being slow, but every request has a deadline: 450 seconds from dispatch, or the `timeout_ms` of a [routing policy](/docs/routing-policies) (1–600 s). A stream still open at the deadline ends with the terminal error event described below. ```bash curl -N https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"claude-sonnet-5","stream":true,"max_tokens":100, "messages":[{"role":"user","content":"Say hello"}]}' ``` Every stream response carries `x-request-id`, `x-tm-provider` (the deployment that served you), `x-tm-attempts`, and `x-tm-upstream-status` headers, plus `x-tm-dropped-params`, `x-tm-upstream-model` and `x-tm-served-by` when they apply (see [Wire compatibility](/docs/compat#response-headers)). ## Frame order — OpenAI surface Each frame is a complete `data: {json}\n\n` line in the documented `chat.completion.chunk` shape. The order is fixed: role-priming delta, content deltas, finish chunk, usage chunk, `data: [DONE]`. ```text data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]} data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]} data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":" there."},"finish_reason":null}]} data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{},"finish_reason":"stop","native_finish_reason":"end_turn"}]} data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":9,"total_tokens":21,"prompt_tokens_details":{"cached_tokens":0,"cache_write_tokens":0},"completion_tokens_details":{"reasoning_tokens":0},"cost":0.000114}} data: [DONE] ``` Notes on this surface: - The **role-priming delta** (`{"role":"assistant","content":""}`) is emitted as soon as the upstream proves alive — it is a liveness signal, not content. - Tool calls stream as `delta.tool_calls` entries with stable ascending `index` values; the arguments arrive as string fragments to concatenate. A tool called with no arguments gets one `"{}"` fragment. - Reasoning models stream their thinking as `delta.reasoning_content`. - Refusal text streams as `delta.refusal`. - On a translated stream the finish chunk also carries `native_finish_reason`, the provider's own stop reason (here `end_turn`). - The usage chunk is always sent, zero-filled if the provider reported nothing, unless you sent `stream_options.include_usage: false`. - `[DONE]` is only ever sent after a successful finish — **never after an error** (see below). ## Frame order — Anthropic surface Named SSE events (`event: \ndata: {json}\n\n`) exactly as the Anthropic Messages API frames them: `message_start`, `content_block_start` / `content_block_delta` / `content_block_stop`, `message_delta`, `message_stop`. ```text event: message_start data: {"type":"message_start","message":{"id":"msg_2f6d8e3b","type":"message","role":"assistant","model":"gpt-4o-mini","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":0,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":0}}} event: content_block_start data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}} event: content_block_delta data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello there."}} event: content_block_stop data: {"type":"content_block_stop","index":0} event: message_delta data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"input_tokens":12,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":9,"cost":0.000007}} event: message_stop data: {"type":"message_stop"} ``` Notes on this surface: - On a translated stream the token counts in `message_start` are zeros; the authoritative usage (and `cost`) arrives in the final `message_delta`. Read usage there, not from `message_start`. On a relayed stream `message_start` carries the provider's own counts. - Tool calls stream as `tool_use` content blocks with `input_json_delta` fragments; thinking streams as `thinking` blocks with `thinking_delta` (and `signature_delta` where the provider signs it). - `input_tokens` excludes cache reads and writes, per the [billing contract](/docs/compat) — cache traffic is reported separately in `cache_read_input_tokens` and `cache_creation_input_tokens`. ## Usage is always in the stream You never opt in to stream usage. The gateway injects `stream_options.include_usage` into every OpenAI-dialect upstream dispatch itself (without it, OpenAI streams carry no usage at all). An explicit `stream_options: {"include_usage": false}` is honored: you are billed the same, but no usage chunk is emitted to you. Every other stream ends with provider-reported token counts, and every **billed** stream carries `usage.cost` (USD) — computed with the exact integer micro-USD math the ledger settles with, covering the full request debit including any attempts that failed over before your answer started. > [!TIP] > `cost` on the wire is the number you are charged. You can recompute it any > time from the token counts and the public prices at > [https://app.routerplus.com/api/models.json](https://app.routerplus.com/api/models.json) — settlement > rounds down, in your favor. ## Keep-alive frames Slow reasoning models can go quiet for a long time mid-answer, and idle TCP connections get killed by proxies, load balancers, and some HTTP clients. So once your stream has started, any silence of 15 seconds or more produces a keep-alive frame: - **OpenAI surface:** an SSE comment — `: processing` — which the SSE spec requires clients to ignore. SDKs never see it. - **Anthropic surface:** a native ping event — `event: ping` / `data: {"type":"ping"}` — the same frame Anthropic's own API sends, which Anthropic SDKs already skip. Keep-alives are measured from the last real write, so the silence you can observe is bounded at ~15s, and nothing — keep-alives included — is ever sent after a terminal event. ## Cancelling a stream Close the connection (abort the HTTP request) and the gateway immediately aborts the upstream request, so the provider stops generating. The billing semantics are deliberate and worth knowing: - **You are billed for what was generated** up to the abort — the attempt is recorded with outcome `cancelled` and the usage streamed so far, visible at `GET /v1/generation?id=`. If the provider reported no usage, the output you received is estimated at about four characters per token. - **A cancellation never counts against the provider's uptime.** You hung up; the provider did nothing wrong. Cancelled attempts are excluded from the health circuits and the public uptime numbers (see [Routing & failover](/docs/routing)). ```typescript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY, }); const stream = await client.chat.completions.create({ model: "claude-sonnet-5", max_tokens: 4096, messages: [{ role: "user", content: "Explain SSE, thoroughly." }], stream: true, }); let chars = 0; for await (const chunk of stream) { chars += chunk.choices[0]?.delta?.content?.length ?? 0; if (chars > 2000) { stream.controller.abort(); // upstream generation stops; billed to here break; } } ``` ## When a provider dies mid-stream If the serving provider fails **after** your answer started, the stream does not silently end and the request is never restarted on another provider (the commit boundary — see [Routing & failover](/docs/routing)). Instead you get one terminal error event inside the already-open 200 stream: **OpenAI surface** — any known usage (and its `cost`) is emitted first, so the billed partial attempt stays recomputable from the wire, then a final chunk with `finish_reason: "error"` and a top-level `error`: ```text data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":""},"finish_reason":"error"}],"error":{"code":"upstream_error","message":"upstream failure","type":"gateway_error","metadata":{"error_type":"upstream_error"}}} ``` **Anthropic surface** — a native `error` event carrying both the Anthropic-native type string and the stable canonical `error_type`: ```text event: error data: {"type":"error","error":{"type":"api_error","message":"upstream failure","error_type":"upstream_error"}} ``` Nothing follows a terminal event — no `[DONE]`, no further frames. The same event ends a stream that reaches the request deadline, and a stream whose provider closed the connection without a proper terminator. > [!NOTE] > The Anthropic wire has no legal usage slot outside `message_delta`, so an > Anthropic-surface stream that dies mid-answer cannot carry its partial usage > in-band. `GET /v1/generation?id=` returns the settled numbers > for exactly this case. This is a documented limitation of the wire format, > not of the ledger. The full taxonomy of `error_type` values and what to do about each lives in [Errors & remediation](/docs/errors). ## SDK streaming examples Both official SDKs work unmodified — swap the base URL and key. ```python import os from openai import OpenAI client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) stream = client.chat.completions.create( model="claude-sonnet-5", max_tokens=200, messages=[{"role": "user", "content": "Write a haiku about failover."}], stream=True, ) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) if chunk.usage: # the final usage chunk has an empty choices list print(f"\ncost: ${chunk.usage.cost}") ``` ```python import os import anthropic client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"]) with client.messages.stream( model="gpt-4o-mini", # yes — a GPT model over the Anthropic wire max_tokens=200, messages=[{"role": "user", "content": "Write a haiku about failover."}], ) as stream: for text in stream.text_stream: print(text, end="", flush=True) final = stream.get_final_message() print(f"\nstop_reason: {final.stop_reason}") ``` --- Source: https://app.routerplus.com/docs/errors.md # Errors Every failure carries one canonical error class — a short stable string like `rate_limit` or `insufficient_quota` — in a fixed spot in the body and in the `x-tm-error-code` response header. The body's native shape follows the surface you called (OpenAI-compatible or Anthropic-compatible), so your SDK's built-in error handling keeps working; the canonical class is there so your code never has to parse prose. A failure that happened after your request was accepted — at a limit, at a provider, or in our infrastructure — also says where it came from. The `x-tm-error-origin` header (and `origin` in the body's metadata) is one of `gateway_admission` (a limit or balance of yours), `upstream_quota` (a provider's declared capacity), `upstream` (the provider answered), `gateway_infrastructure` (our side) or `authorization` (your key or connection). When a limit refused the request, `x-tm-limit-scope`, `x-tm-limit-kind` and, where known, `x-tm-limit-id` name it. See [Limits and capacity](/docs/admission). ## Where the class lives OpenAI surface (`POST /v1/chat/completions`): ```json { "error": { "code": "model_unavailable", "message": "model \"gpt-5-nano\" is not in the catalog; GET /v1/models lists what this key can serve", "type": "invalid_request_error", "metadata": { "error_type": "model_unavailable" } }, "request_id": "6f8f57b2-..." } ``` Anthropic surface (`POST /v1/messages`): ```json { "type": "error", "error": { "type": "not_found_error", "message": "model \"gpt-5-nano\" is not in the catalog; GET /v1/models lists what this key can serve", "error_type": "model_unavailable" }, "request_id": "6f8f57b2-..." } ``` Four rules govern these bodies: - `error.metadata.error_type` (OpenAI surface) or `error.error_type` (Anthropic surface) is **always the canonical class**. Every error body is generated by the gateway; a provider's own error body is never relayed. The provider's HTTP status is in `x-tm-upstream-status` when it answered. - On the Anthropic surface, `error.type` additionally maps to the **native Anthropic type string** (`rate_limit_error`, `billing_error`, ...) so the Anthropic SDK's own error classification and retry behavior work unmodified. - On the OpenAI surface, `error.code` carries the canonical class. Switch on `error.code`, `metadata.error_type`, or the HTTP status — not on `error.type`, which is a generic string. The one exception is out-of-credits: there both `code` and `type` are `insufficient_quota`, byte-matching OpenAI's own wire so OpenAI SDKs recognize it natively. - Errors raised before routing (bad key, unknown route, a crash in our handler) pick the shape from the path: requests to `/v1/messages` get Anthropic-shaped bodies, everything else OpenAI-shaped. The metadata carries a few more fields after a limit or a balance refused the request: `origin`, `limit_scope`, `limit_kind`, `limit_id`, `retryable`, and `retry_at` or `reset_at` when one is known. On the OpenAI surface they sit beside `error_type` in `error.metadata`; on the Anthropic surface in `error.metadata`. ## The canonical classes | `error_type` | HTTP | OpenAI `error.code` | Anthropic `error.type` | Meaning | |---|---|---|---|---| | `auth` | 401 (or a relayed 401/402/403) | `auth` | `authentication_error` | your key is missing, wrong, or disabled — or, rarely, the marketplace's own provider account failed on every candidate (see below) | | `invalid_request` | 400 | `invalid_request` | `invalid_request_error` | body is not valid JSON, `n>1`, or a request no deployment can serve as sent — `temperature > 1`, `response_format` or `thinking` across a dialect boundary, unsupported content (images, audio) across a dialect boundary, or a content part the model does not take (an image, file, audio or video part that no route of the model declares) — the message names the field; a request-level credential or destination (`api_key`, `base_url`, `connection_id`, …) or an unknown `provider` control; on `/v1/images/generations` and `/v1/videos`: a value outside the model's accepted values (the message names the field), `n` outside 1–4, `stream`/`partial_images`, `response_format: url`, an image input on the videos route, a model on the wrong route — and an image or video model on a chat route (the message names the right route); on `/v1/decisions`: a question in another decision model's format (`schema` on Jev, Mercury Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider v1 27B or Sage, `questions` on GLiNER-2.5-Decide), a question kind that does not fit the state, a field the model does not take, two GLiNER results with one name, or, on Perplexity Decider v1 27B, an object with `"type": "image_url"` in the state or a question, because decision models take text and JSON (the message names the field) | | `model_not_priced` | 400 | `model_not_priced` | `invalid_request_error` | the model is listed but has no price row; we refuse rather than guess $0 | | `context_overflow` | provider's status (usually 400; 413 from Cloudflare on Clef and Clef-flash) | `context_overflow` | `invalid_request_error` | prompt exceeds the model's context window; on `/v1/decisions`, Cloudflare refuses a Clef or Clef-flash request that it estimates above 65,536 tokens (its 413, code 5021), and Perplexity refuses a Perplexity Decider v1 27B question whose prompt (the state and that question) is above 262,144 tokens (its 400); the message carries the provider's words | | `content_policy` | provider's status (usually 400) | `content_policy` | `invalid_request_error` | the provider refused the content; an OpenAI image moderation refusal (`moderation_blocked`) is relayed at the provider's own status (usually 400): never billed, never rerouted, never a health strike | | `model_unavailable` | 404 / 502 / 503 | `model_unavailable` | `not_found_error` | 404: not in the catalog (message echoes your requested id), or not served by your key's connection · 502: connection to the provider failed on every candidate · 503, on `/v1/decisions`: the provider refuses our account for capacity: Levanto's monthly allowance for Sage is used up (the message says Sage is out of capacity until the provider's next billing period), or Fastino refuses GLiNER-2.5-Decide for credit or billing, or Bespoke Labs refuses Bespoke Nimble v3 for credit (its 402), or Cloudflare refuses Clef or Clef-flash for our token or plan (its 401 or 403) or because our account's free daily allocation is used up (its 429 with code 3036); a 401, 402 or 403 from any System One provider on our key reads the same (TypeSafe or OpenRouter, on Jev or Mercury Decide, and RouterPlus, on Decider 2B and Kev 4B, too, and Perplexity, on Perplexity Decider v1 27B), and on a connection with your own key it keeps its status — the message says the model is out of capacity at the provider | | `upstream_error` (timeout) | 504 | `upstream_error` | `api_error` | the provider exceeded the gateway's wait: 20 s to first response headers on streams, the generation budget (450 s) on non-stream calls, 60 s to accept a video job, or 60 s for an answer on `/v1/decisions` (a GLiNER-2.5-Decide cold start can outlast it; a RouterPlus cold start, 15 to 20 s, fits inside it). On `/v1/decisions` a System One provider's own 408 (Cloudflare's timeout on Clef and Clef-flash) is this class too, and so is Perplexity's own 504 on Perplexity Decider v1 27B (it counts against the circuit, as every 5xx does). Retryable; a lone timeout never trips the health circuit | | `not_found` | 404 | `not_found` | `not_found_error` | no such route — or, on `/v1/videos/{id}`, no such video for your organization, or on `/v1/generation`, an id that is not a request id | | `request_too_large` | 413 | `request_too_large` | `request_too_large` | body over 10 MB (1 MB on `POST /v1/videos` and `POST /v1/decisions`) | | `rate_limit` | 429 | `rate_limit` | `rate_limit_error` | a rate limit refused the request: the key's RPM, one of the organization's shared limits, a provider pool's declared capacity or its cooldown after a provider 429, or the provider rate-limited every candidate. `x-tm-limit-scope` and `x-tm-limit-kind` say which, and `x-tm-limit-id` whether a pool limit was the pool's (`pool:…`) or your organization's share of it (`pool-share:…`). On `/v1/decisions`, Mercury Decide's pool is small (20 requests a minute in all, 15 per organization; 5 in flight, 3 per organization) and OpenRouter caps free requests per day across our account: both come back as this class, and a provider's 429 never opens the deployment's health circuit; Bespoke Nimble v3's pool takes 7 requests at once, 5 per organization, because Bespoke allows 8 at once for our whole account and one is kept for our test environment, and Bespoke's own 429 comes back as this class too; Clef and Clef-flash share one pool (200 requests a minute, 150 per organization; at most 8 at a time, 6 per organization), and Cloudflare's 429 when Workers AI is busy (code 3040) comes back as this class too, while its code 3036 is the 503 `model_unavailable` above; Decider 2B and Kev 4B share one RouterPlus pool (11,400 requests a minute, 59 in flight, 75 % per organization), so your organization's own limits bind first, and RouterPlus's own 429 (about 200 requests a second for the whole endpoint) comes back as this class too and pauses the pool for all three; Perplexity Decider v1 27B's pool takes 540 requests a minute (405 per organization) and 8 at once (6 per organization), its hold counts the state once per question, so a long state with many questions can pass a token limit on its own (`x-tm-limit-kind: tpm`: ask fewer questions per call), and Perplexity's own 429 (10 requests a second for our organization) comes back as this class too | | `email_not_verified` | 403 | `email_not_verified` | `permission_error` | the key's account signed up but has not opened its email verify link yet; every request is refused until it does, even one that would cost nothing | | `insufficient_quota` | 429 | `insufficient_quota` | `billing_error` | balance too low, or a monthly spend cap hit | | `usage_limited` | 429 | `usage_limited` | `rate_limit_error` | unusual usage was detected on an organization that has only trial credit, and this request was limited; `retry-after` says when to try again | | `usage_restricted` | 403 | `usage_restricted` | `permission_error` | unusual usage was detected on an organization that has only trial credit, and this kind of request is refused until support reviews the account | | `upstream_error` | 5xx, or provider's 4xx | `upstream_error` | `api_error` | provider failed and no alternate could serve; at 4xx the request itself was rejected; on the images route a 200 that is unparseable or carries no image is a 502 that bills nothing | | `gateway_error` | 400 / 500 / 502 / 503 / 504 | `gateway_error` | `api_error` | 400: an output bound the gateway refused — `max_tokens` outside 1–32,768, or `max_tokens` and `max_completion_tokens` both present and different (`x-tm-limit-kind: output_bound`) · 500: our bug · 502: the provider's image response exceeded the 32 MiB cap, or a 2xx our translator cannot represent; nothing billed · 503: no capacity right now — every deployment serving the model is cooling down (`x-tm-limit-kind: health`, `retry-after: 5`), the platform or this worker is at capacity, the shared limit store or the ledger journal is unavailable; nothing was sent to a provider and nothing was charged · 504: the request's routing deadline (450 s, or the policy's `timeout_ms`) passed before a provider could be tried | When the status comes from a provider, the provider's own HTTP status is also exposed in the `x-tm-upstream-status` header. ## What to do, per class | Class | Retry? | Remediation | |---|---|---| | `auth` (401 from us) | no | send the key as `Authorization: Bearer tm_vk_...` or `x-api-key: tm_vk_...` — both work on every endpoint; complete browser signup and create a key in the console at [https://app.routerplus.com/console/keys](https://app.routerplus.com/console/keys) | | `auth` (relayed 401/402/403) | no | nothing — this is our provider account, not yours; see the note below | | `invalid_request` | never blind-retry | fix what the message names, then resend | | `model_not_priced` | no | report it with the `x-request-id` — a listed model without a price is our data bug | | `context_overflow` | never blind-retry | shorten the input or pick a larger-context model; deliberately **not** failed over | | `content_policy` | never blind-retry | change the prompt; **never rerouted** to another provider | | `model_unavailable` (404) | no | pick an id from `GET /v1/models` (or the public list at `https://app.routerplus.com/api/models.json`) | | `model_unavailable` (502) | yes | one retry after ~2 s; every candidate's connection failed | | `model_unavailable` (503, decisions) | not soon | the provider's capacity comes back when its allowance or credit does (for Sage, at Levanto's next billing period; for Clef and Clef-flash after Cloudflare's code 3036, at 00:00 UTC); use another decision model meanwhile, in its own question format. `metadata.retryable` is `false`. While the provider refuses, most requests to that model get `gateway_error` (503, "all deployments cooling down") instead, because each refusal counts against the provider's circuit: treat that the same way on Sage, GLiNER-2.5-Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B and Perplexity Decider v1 27B (Cloudflare's code 3036 is a 429, and a 429 never counts) | | `not_found` | no | endpoints: `/v1/chat/completions`, `/v1/messages`, `/v1/messages/count_tokens`, `/v1/images/generations` (alias `/images/generations`), `/v1/videos`, `/v1/videos/{id}`, `/v1/videos/{id}/content`, `/v1/decisions`, `/v1/models`, `/v1/usage`, `/v1/generation`, `/v1/limits`, `/v1/route` | | `request_too_large` | no | trim the body under 10 MB (1 MB for a video or decisions request) — the images route also caps the provider response at 32 MiB, but that is a 502 `gateway_error`, not a 413 | | `rate_limit` | yes | wait `retry-after`, retry once — details in [Rate limits & spend caps](/docs/limits) and [Limits and capacity](/docs/admission) | | `email_not_verified` | **no** | open the verify link in the signup email; the key works within seconds of the click | | `usage_limited` | after `retry-after` | honor `retry-after`; if it keeps happening, contact support | | `usage_restricted` | **no** | contact support to have the account reviewed | | `insufficient_quota` | **no** | retrying cannot help: add credits, or raise the cap at [https://app.routerplus.com/console/billing](https://app.routerplus.com/console/billing); on cap breaches the exact reset time is in the message and `x-tm-cap-reset` | | `upstream_error` (5xx) | yes | one retry after ~2 s; a failover already happened if one was possible | | `upstream_error` (4xx) | never blind-retry | the request itself was rejected — the message names the provider and its status; `x-tm-upstream-status` carries it | | `gateway_error` (400) | never blind-retry | send one output limit, between 1 and 32,768 | | `gateway_error` (500) | no | report with the `x-request-id` | | `gateway_error` (502, images) | yes, once | nothing was billed; lower `n` or the quality | | `gateway_error` (503) | yes | wait `retry-after` and resend — the request never reached a provider, so a retry cannot double-bill | | `gateway_error` (504) | yes | resend; the deadline passed while candidates were being tried | ## Failure headers | Header | When | Meaning | |---|---|---| | `x-request-id` | always | quote it in any report; also the id for `GET /v1/generation?id=` | | `x-tm-error-code` | every failure | the canonical class | | `x-tm-error-origin` | limit, provider and infrastructure failures | `gateway_admission`, `upstream_quota`, `upstream`, `gateway_infrastructure` or `authorization`; a request refused for its own shape (bad JSON, `n>1`, an unknown model) carries none | | `x-tm-limit-scope`, `x-tm-limit-kind`, `x-tm-limit-id` | a refused limit | scope (`platform`, `org`, `workspace`, `principal`, `key`, `model`, `pool`, `endpoint`), the kind of limit (`rpm`, `rps`, `rps_queue`, `tpm`, `concurrency`, `spend`, `health`, …), and the limit's id when one exists | | `x-tm-attempts` | after ≥1 dispatch | physical provider attempts, failovers included | | `x-tm-upstream-status` | when a provider answered | the provider's own HTTP status | | `retry-after` | 429 / 503 | seconds to wait — a rate limit says when its window frees; provider-sent values are honored but capped at 60; the all-cooling-down 503 sends 5; the ledger-journal 503 sends 2 | | `x-tm-cap-reset` | spend-cap 429 only | exact ISO instant the cap resets (first instant of the next UTC month) | `x-tm-error-code` is not CORS-exposed; browser code reads the class from the body. ## Strict by design Five refusals people trip on are deliberate — the alternative in each case is billing you for something wrong. **`n>1` is a typed 400, not a degraded answer.** The normalizer emits exactly one choice; accepting `n: 4` and silently billing a garbled single-choice response would be dishonest. ```bash curl -s https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"claude-haiku-4-5","n":4,"messages":[{"role":"user","content":"hi"}]}' ``` ```json { "error": { "code": "invalid_request", "message": "n>1 is not supported in v1; request a single choice", "type": "invalid_request_error", "metadata": { "error_type": "invalid_request" } }, "request_id": "..." } ``` **Content a model does not take is a typed 400, not an answer that ignored it.** A provider that does not read images can drop the image, answer the rest, and still bill the request. So an image, file, audio or video part goes only to a route whose listing takes that kind, and a part that no route of the model takes is refused before anything is reserved or sent. `GET /v1/models` lists what each model takes in `architecture.input_modalities`. ```bash curl -s https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"deepseek/deepseek-v4-flash","messages":[{"role":"user","content":[{"type":"text","text":"What is in this picture?"},{"type":"image_url","image_url":{"url":"https://example.com/cat.png"}}]}]}' ``` ```json { "error": { "code": "invalid_request", "message": "messages[0].content[1]: model \"deepseek/deepseek-v4-flash\" does not take image input; it takes text", "type": "invalid_request_error", "metadata": { "error_type": "invalid_request" } }, "request_id": "..." } ``` **No silent model substitution.** An unknown model id is a 404 whose message echoes exactly the string you sent — we never quietly swap in a "close enough" model. The 404 is identical whether the id never existed or simply is not served, so the error is not an existence oracle. **A content-policy refusal never reroutes.** A refusal is not shopped around to a more permissive provider — that is a hard routing rule with no exceptions. If you get `content_policy`, the fix is the prompt. **Images are all-or-nothing.** A failed, over-cap or malformed generation is a 502 and bills nothing; `n` above 4 on the images route is a typed 400 like `n > 1` on chat. See [POST /v1/images/generations](/docs/api-images). ## Failover, and what the error you see means Before any output has reached you, the gateway fails over between deployments invisibly, at zero backoff. Classes that are failover-eligible: `rate_limit`, `auth` (upstream), `model_unavailable`, and 5xx `upstream_error`. Classes that never fail over: `content_policy` (never, as above), `context_overflow` (its own typed class — reshape the request instead), and provider 4xx rejections (the request is the problem). So if a failover-eligible class reaches you, every healthy candidate was tried — `x-tm-attempts` says how many. > [!NOTE] > A relayed 402 (class `auth`) is never your fault. An upstream 401/402/403 means the **marketplace's** account with that provider has a credential or payment problem. Payability is treated as a health signal: the error is failover-eligible and counts against that provider's uptime, so an unpayable provider drains traffic automatically. You would only ever see it when every deployment serving the model has the same problem. On `/v1/decisions` most of these refusals come back as 503 `model_unavailable` instead, with `metadata.retryable: false` (see the table above). ## Mid-stream failures (HTTP 200 already committed) Once your answer has started streaming, the gateway never switches providers — a provider death mid-answer becomes exactly one terminal error event **inside** the 200 stream, and nothing follows it (no `data: [DONE]`, no keep-alives, no more chunks). OpenAI surface — a real chunk envelope with `finish_reason: "error"` and a top-level `error`, so SDKs raise and naive parsers still terminate cleanly: ```json { "id": "chatcmpl-...", "object": "chat.completion.chunk", "created": 1757000000, "model": "claude-haiku-4-5", "choices": [{ "index": 0, "delta": { "content": "" }, "finish_reason": "error" }], "error": { "code": "upstream_error", "message": "upstream failure", "type": "gateway_error", "metadata": { "error_type": "upstream_error" } } } ``` Anthropic surface — a native `error` event, with the canonical class riding alongside the native type: ```json { "type": "error", "error": { "type": "api_error", "message": "upstream failure", "error_type": "upstream_error" } } ``` On the OpenAI surface, any known usage and cost are emitted before the terminal error, so the wire stays billable-recomputable. The Anthropic wire has no legal usage slot outside `message_delta`; for an Anthropic-surface stream that died mid-answer, recompute from `GET /v1/generation?id=`. Retrying after a mid-stream error is your call — the answer never restarts on a different provider by itself. ## Reading the class in code ```python import os from openai import OpenAI, APIStatusError client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) try: out = client.chat.completions.create( model="claude-haiku-4-5", messages=[{"role": "user", "content": "hello"}], ) except APIStatusError as e: err = e.response.json().get("error") or {} kind = (err.get("metadata") or {}).get("error_type") \ or e.response.headers.get("x-tm-error-code") origin = e.response.headers.get("x-tm-error-origin") request_id = e.response.headers.get("x-request-id") # kind is one of the canonical classes in the table above ``` ## Retry discipline for agents - `rate_limit` → wait `retry-after`, retry once. - `email_not_verified` → do **not** loop; ask the human to open the verify link in their email. - `insufficient_quota` → do **not** loop; surface to a human (money). See [Rate limits & spend caps](/docs/limits). - `usage_limited` → do **not** loop; honor `retry-after` once, then tell the human to contact support. - `usage_restricted` → do **not** retry; tell the human to contact support. - `upstream_error` at 5xx → one retry after ~2 s, then surface with the `x-request-id`. At 4xx the request itself is the problem — never blind-retry. - `upstream_error` or `gateway_error` at 502 on the images route → one retry; nothing was billed. - `gateway_error` at 503 or 504 → wait `retry-after` (or ~2 s), retry; nothing was billed. - `model_unavailable` at 404 → pick a real id from `GET /v1/models`; at 502 one retry. - `content_policy`, `context_overflow`, `invalid_request`, `gateway_error` at 400 → never blind-retry; the request must change first. ## A video that fails A video job ends `failed` without an HTTP error: `POST /v1/videos` answered 200 when the job was accepted, and the failure arrives later, on `GET /v1/videos/{id}`, as `status: "failed"` with `error: {code, message}`. The code is one of `video_generation_failed`, `content_policy_violation`, `video_cancelled`, `video_expired` or `video_timeout`. A failed job costs nothing; `content_policy_violation` means change the prompt. See [POST /v1/videos](/docs/api-videos). --- Source: https://app.routerplus.com/docs/routing.md # Routing & failover You ask for a model; the gateway picks a deployment that serves it and fails over between deployments when providers misbehave. Three guarantees shape everything on this page: 1. **The model you name is the model you get.** An unknown model is an honest 404 naming the id — never a silent substitute. 2. **Failover only happens before your answer starts.** Once the first token of real output reaches you, the request is committed to that provider forever. 3. **A content-policy refusal is never rerouted.** Shopping a refused prompt around providers is a policy decision we refuse to make for you. ## How a request picks a provider The catalog is a list of **deployments** — each one a provider endpoint with a wire dialect, the model ids it serves, and a routing `priority`. Today most models have two: the lab's own API first, then OpenRouter as the fallback (`anthropic` then `openrouter` for Claude models, `openai` then `openrouter` for GPT models). Selection is a filter pipeline, entirely in-memory — routing itself makes no remote calls. (Billed requests separately consult the ledger and the shared admission store for authentication, limits and the reserve check.) 1. **Model filter** — keep deployments whose model list contains the requested id, exactly. No fuzzy matching, no aliases. The one resolution: a dated snapshot of a listed id (`claude-haiku-4-5-20251001`) routes and bills as its family, and your original id still goes to the provider verbatim. 2. **Health filter** — drop deployments whose circuit is open (cooling down). 3. **Priority sort** — survivors are stable-sorted by ascending `priority`; ties keep catalog order. Dispatch then walks that ordered list. The first deployment to produce real output wins; failover-eligible failures move to the next candidate with zero backoff. Which dialect a deployment speaks is invisible to you — the gateway translates requests, streams, and errors both ways, so any listed model is callable from either surface. A key with a BYOK connection or a routing policy runs the same pipeline over its own routes, with the policy's order, funding rule and attempt limit on top; `POST /v1/route` explains the plan for a request without dispatching it. See [Routing policies](/docs/routing-policies). A [dedicated endpoint](/docs/dedicated-endpoints) (`/`) routes to capacity reserved for your organization first, then to public deployments of the same model that meet the endpoint's constraints. It never routes to a different model. A [closed](/docs/dedicated-endpoints#closed-endpoints) endpoint names RouterPlus as the provider, whichever route served the request. ## The commit boundary Failover is governed by one line: **the request commits to a deployment at its first semantic output** — the first text delta, reasoning delta, tool call, refusal, or usage report that reaches you. **Before commit**, failover is invisible. Failover is possible only while nothing has been written to you. A provider that fails before its stream produces a first frame — connection refused, a 5xx, a death before any bytes, or blowing the 20-second headers deadline — is skipped and the next candidate is tried immediately; your response headers are sent only once a live attempt starts writing. Once the headers and (on the OpenAI surface) the role-priming delta have gone out, the gateway no longer switches providers: a failure after that point becomes the terminal in-stream error described in [Streaming](/docs/streaming), even if no semantic output was produced. **After commit**, the gateway will never switch providers. Two different models do not produce interchangeable halves of one answer, and splicing them silently would be a lie about what you received. If the provider dies mid-answer you get the tokens it produced plus one terminal error event inside the 200 stream (exact frames in [Streaming](/docs/streaming)) — retrying is your decision, and the retry starts fresh. > [!NOTE] > You pay for every settled attempt, including ones discarded by pre-commit > failover — a provider that dies after reporting prompt usage still billed us > for that prompt, and pass-through pricing passes it through. The inline > `usage.cost` on your response is the **full request debit** across all > attempts, so the wire number always matches the ledger. `x-tm-attempts` tells > you how many physical dispatches happened. ## No silent model substitution If the requested model is not in the catalog, the response is a 404 that names it: ```json { "error": { "code": "model_unavailable", "message": "model \"gpt-5-ultra\" is not in the catalog; GET /v1/models lists what this key can serve", "type": "invalid_request_error", "metadata": { "error_type": "model_unavailable" } }, "request_id": "…" } ``` There is no "closest match" fallback and no default model. The same honesty applies to requests we can't serve faithfully: `n>1` is a typed 400 (the gateway normalizes every stream to a single choice, and billing you for a garbled multi-choice response would be worse than refusing). ## What never fails over Failover eligibility is decided once, at error-normalization time — routing never inspects provider-raw errors. Two classes are hard-excluded: - **`content_policy`** — checked before every status-code rule, so no HTTP status can override it: a refusal never retries and **never** reroutes to another provider. Error metadata on this path never echoes your flagged input. If you want a second opinion from a different provider, that is an explicit new request you make yourself. - **`context_overflow`** — its own typed class. Another deployment of the same model has the same context window; the fix is reshaping the request or choosing a larger-context model, not blind rerouting. Other buyer-fault 4xx errors (malformed request, etc.) also return directly: retrying an invalid request elsewhere just spends your money on the same error. ## Failover eligibility by error class | Upstream signal | `error_type` | Retryable | Fails over | |---|---|---|---| | 429 (Retry-After honored, capped 60s) | `rate_limit` | yes | yes | | 401 / 402 / 403 from the provider | `auth` | no | **yes** | | Context / length errors | `context_overflow` | no | no | | Content policy / refusal shapes | `content_policy` | no | **never** | | 404 / model_not_found upstream | `model_unavailable` | no | yes | | 5xx / 529 / overloaded | `upstream_error` | yes | yes | | Other 4xx | `upstream_error` | no | no | > [!NOTE] > A 401/402/403 **from a provider** is our account problem with that provider — > not yours. It routes around the deployment and counts against that provider's > health, so a provider we can no longer pay drains traffic automatically. You > only see an `auth` error if every candidate failed. ## Health circuits and cooldowns Every deployment has an independent circuit breaker: - **Two consecutive counted failures open the circuit** for a 30-second cooldown. While open, the deployment is skipped by routing. - **After the cooldown the circuit goes half-open**: exactly one live request is admitted as a probe while everyone else keeps skipping. A successful probe closes the circuit; a failed probe re-opens it with a fresh cooldown. This avoids the stampede of a blind TTL expiry re-admitting all traffic at once. - **Only failover-eligible failures charge the circuit.** Buyer-fault errors (400/413-class, content policy, context overflow) never do — your malformed request must not take a healthy provider out of rotation for everyone else. (An upstream 429 does charge the circuit — it is failover-eligible — but is excluded from uptime scoring as buyer-caused. It also cools that provider's admission pool for the provider's `retry-after`, 1 to 60 seconds.) A client cancellation likewise counts in the provider's favor: you hung up; it did nothing wrong. The same fairness rule shapes the public numbers: buyer-caused failures and cancellations are excluded from uptime scoring, and a model page shows no uptime percentage for a provider we call directly until it has 100 counted requests in the window — a 3-for-3 "100%" would be noise dressed up as a guarantee. ## When everything is cooling down If every deployment serving your model has an open circuit, the request is refused at once rather than queued: a `503` with `retry-after: 5`, `x-tm-error-code: gateway_error`, `x-tm-limit-kind: health` and the message `all deployments cooling down`. It is not a 404 — the model exists — and it is transient: a cooldown lasts 30 seconds, after which one request is let through as the probe. Retry after a few seconds. ## Reading what happened Every response tells you how it was served: | Header | Meaning | |---|---| | `x-request-id` | your handle for audit and support (always present) | | `x-tm-provider` | the deployment that served the request; `routerplus` on a closed dedicated endpoint | | `x-tm-served-by` | on an aggregator route, the provider that actually ran the model; absent on a closed dedicated endpoint | | `x-tm-upstream-model` | the id sent to the provider, when it differs from the one you sent; absent on a closed dedicated endpoint | | `x-tm-attempts` | physical dispatches, failovers included | | `x-tm-upstream-status` | the HTTP status the provider actually returned | | `x-tm-dropped-params` | request fields the gateway stripped before dispatch, named | | `x-tm-route-plan-id` | the plan id, for correlation with `/v1/route` | | `x-tm-dedicated-endpoint`, `x-tm-route-role` | on a dedicated endpoint: its id, and whether your dedicated capacity (`primary`) or the shared pool (`fallback`) served | | `x-tm-queue-ms` | on a dedicated endpoint with a [paced limit](/docs/dedicated-endpoints#the-paced-limit): how long the request waited for its turn, in milliseconds | | `x-tm-remaining-rpm`, `x-tm-remaining-tpm` | what your rate limits had left at admission | | `x-tm-error-code` | canonical error class, on failures | | `x-tm-error-origin` | on failures: `upstream`, `gateway_admission`, `upstream_quota`, `gateway_infrastructure` or `authorization` | | `retry-after` | present on rate limits and cooldowns; capped at 60s | On failures the body carries the stable canonical class — `error.metadata.error_type` on the OpenAI surface, `error.error_type` on the Anthropic surface — alongside each dialect's native type string so your SDK's built-in retry/backoff classification keeps working. The body is always the gateway's own shape; a provider's raw error body is never relayed. The full table is in [Errors & remediation](/docs/errors). For the complete per-attempt story, the generation audit endpoint returns every physical attempt for a request id — which deployments were tried, in what order, and what each one cost: ```bash curl -s -H "Authorization: Bearer $TM_API_KEY" \ "https://api.routerplus.com/v1/generation?id=" ``` A Claude request whose first attempt failed over from the lab's API to OpenRouter (trimmed: each attempt also carries its billing source, price snapshot, error origin, admission and route context, and dropped parameters): ```json { "request_id": "2f6d8e3b-…", "admission_events": [], "attempts": [ { "attempt": 1, "deployment": "anthropic", "model": "claude-sonnet-5", "outcome": "failed", "usage_provenance": "unknown", "tokens": { "input": 0, "cached": 0, "cache_write": 0, "output": 0, "reasoning": 0 }, "reserved_max_usd": 0.04, "cost_usd": 0, "error_code": "upstream_error", "dispatched_at": "2026-09-04T09:12:44.120Z" }, { "attempt": 2, "deployment": "openrouter", "model": "claude-sonnet-5", "outcome": "completed", "usage_provenance": "observed", "tokens": { "input": 412, "cached": 0, "cache_write": 0, "output": 638, "reasoning": 0 }, "reserved_max_usd": 0.04, "cost_usd": 0.007204, "error_code": "", "dispatched_at": "2026-09-04T09:12:45.010Z" } ] } ``` Metadata only — prompts and responses are never stored. `outcome` is one of `completed`, `failed`, `cancelled`, `incomplete` (died after your answer started), or `unknown_after_crash` (settled at the reserved maximum until reconciled). `admission_events` lists refusals that happened before any attempt, such as a rate limit. This is also the recomputation path for the one wire gap described in [Streaming](/docs/streaming): an Anthropic-surface stream that died mid-answer. --- Source: https://app.routerplus.com/docs/limits.md # Rate limits & spend caps Several independent brakes can stop a request before it reaches a provider: the key's requests-per-minute limit, the organization's token and concurrency limits, the provider pool's capacity, the monthly spend caps, and the balance itself. Each refuses with a typed error that names its class in the body and in `x-tm-error-code` — nothing is throttled or queued silently, and a refused request costs nothing. The error shapes follow the same conventions as everything else — see [Errors](/docs/errors). Pools, fairness and what happens when the shared store is down are on [Limits and capacity](/docs/admission). ## Per-key requests per minute Every key has an RPM limit, enforced over a rolling 60-second window per key. The window lives in a shared store, so every gateway instance counts the same requests, and a request's charge can stay in the window for up to 61 seconds. | Key | RPM | |---|---| | First trial key shown after verified browser sign-in | 20 | | Key created in the console or with the identity API | 300 | | The console's managed Playground key | 60 | Until the organization has bought credit (or holds credit we granted), every customer key and the organization as a whole run at the trial rate: 20 requests per minute, however many keys it creates. The first purchase lifts it within seconds; no key changes. Requests served only by your own provider keys (BYOK) and the managed Playground key are not held to it. Historical card-verification bonuses remain promotional credit, not purchases. An organization on trial credit can also be limited when we detect unusual usage. Such a request gets `429 usage_limited` (with `retry-after`) or `403 usage_restricted`, and the message asks you to contact support, who can review the account. A purchase usually lifts it. Over the limit, the request is refused before any reservation or provider work: ``` HTTP/1.1 429 Too Many Requests retry-after: 1 x-tm-error-code: rate_limit x-tm-error-origin: gateway_admission x-tm-limit-scope: key x-tm-limit-kind: rpm x-tm-limit-id: key: ``` ```json { "error": { "code": "rate_limit", "message": "configured quota exhausted", "type": "invalid_request_error", "metadata": { "error_type": "rate_limit", "origin": "gateway_admission", "limit_scope": "key", "limit_kind": "rpm", "limit_id": "key:", "retryable": true, "retry_at": "2026-09-24T10:14:04.000Z" } }, "request_id": "..." } ``` On the Anthropic surface the same refusal is a native `rate_limit_error` carrying the same `error_type` and `metadata`, so the Anthropic SDK's own backoff logic engages. `retry-after` on a rolling window is advisory: the window frees one second at a time, and other requests may take the room first. Honor it and retry; don't hammer. > [!NOTE] > The RPM values are fixed today: there is no self-serve way to raise a key's ceiling. On [`/console/limits`](https://app.routerplus.com/console/limits) you can only lower a key's RPM, and set its other limits. The same class covers a different case: a provider rate-limiting us upstream. That variant fails over to another deployment automatically and cools that provider's pool; you only see a provider-originated `rate_limit` (`x-tm-error-origin: upstream`) when every candidate was throttled, and its `retry-after` honors the provider's value capped at 60 seconds. ## Organization tokens and concurrency Beyond RPM, every organization and every key has a token-per-minute limit and a concurrency limit, enforced in the same shared window: | Scope | RPM | Tokens per minute | Concurrent requests | |---|---|---|---| | Organization (all keys together) | 300 (20 until the first purchase) | 1,000,000 | 8 | | Key | its ceiling above | 1,000,000 | 8 | These are the defaults. Set your own on [`/console/limits`](https://app.routerplus.com/console/limits) or with `POST /api/admission-limits`; a key's RPM cannot go above its ceiling, and a model-scoped limit can be added on top. Under a contract we can replace these defaults, and the keys' RPM ceiling, for your organization; a limit you set yourself then still applies and can only lower them. Tokens are counted as an estimate at admission — the request's UTF-8 bytes plus a media allowance, plus the output bound — and replaced by observed usage when the request completes. A refusal is the same 429 shape as above with `limit_scope` `org` or `key` and `limit_kind` `tpm` or `concurrency`. Every admitted response carries `x-tm-remaining-rpm` and `x-tm-remaining-tpm`: the smallest remaining amount across the scopes it claimed, at that moment. `GET /v1/limits?model=` on the gateway returns the effective limits for your key. Provider capacity is a further scope: each house deployment has a declared pool, and one organization may use at most half of a pool's usable capacity (three quarters on the decision-model pools). Exhausting it is a 429 with `x-tm-error-origin: upstream_quota` and `x-tm-limit-scope: pool`. Details, including workspace and principal scopes, on [Limits and capacity](/docs/admission). A [dedicated endpoint](/docs/dedicated-endpoints) has its own limits, and they replace the defaults above for its traffic. They can include a requests-per-second limit: a count per second, or a [paced limit](/docs/dedicated-endpoints#the-paced-limit) that holds a short burst in a queue. A paced request that waited carries `x-tm-queue-ms`, the wait in milliseconds. Above the hard rate the refusal is `x-tm-limit-kind: rps`; when the wait would be longer than its bound, `rps_queue`. Its refusals carry `x-tm-limit-scope: endpoint`. ## Monthly spend caps Two optional hard caps, set in the console: - an **org-wide** cap covering every key, at [https://app.routerplus.com/console/billing](https://app.routerplus.com/console/billing), and - a **per-key** cap on any individual key, with **Set cap** at [https://app.routerplus.com/console/keys](https://app.routerplus.com/console/keys). Both run on the **UTC calendar month** and reset at the first instant of the next month. They overlap freely; the strictest one wins. In the console, leaving a cap field blank changes nothing and setting it to 0 removes the cap. Keys near their cap are flagged, and per-key month-to-date spend is shown against the cap. ### How enforcement works — reserve-aware, refuse-early Before dispatching anything, the gateway reserves a worst-case cost for the request (see the math below). The cap check then asks: would *settled spend this month + the worst-case cost of everything currently in flight, this request included* exceed the cap? If yes, the request is refused before it costs anything. The in-flight side is claimed synchronously, so concurrent requests cannot race past the cap together. The consequence to design around: the check is conservative. A request near the boundary can be refused even though its eventual settled cost would have squeezed under — the cap refuses early rather than overshoot. A tight `max_tokens` shrinks the worst-case reserve and buys back headroom near the cap. ### The refusal A cap breach is the standard out-of-credits 429 (`insufficient_quota` — the class OpenAI SDKs already recognize natively), with the exact reset instant in both the message and a header: ``` HTTP/1.1 429 Too Many Requests x-tm-error-code: insufficient_quota x-tm-cap-reset: 2026-10-01T00:00:00.000Z ``` ```json { "error": { "code": "insufficient_quota", "type": "insufficient_quota", "message": "org monthly cap exceeded; resets 2026-10-01T00:00:00.000Z", "metadata": { "error_type": "insufficient_quota", "origin": "gateway_admission", "retryable": true } }, "request_id": "..." } ``` The message names the scope that tripped (`org` or `key`). On the Anthropic surface the body is a native `billing_error` carrying the same `error_type` and message. Do not retry before the reset time — no amount of backoff changes a hard cap, whatever `retryable` says. ## The worst-case reserve The number the cap and balance checks use is deliberately pessimistic: | Input | Value | |---|---| | Estimated input tokens | one per UTF-8 byte of the request body, plus 65,536 per image or document part; an image sent inline as base64 counts the 65,536 only, not its bytes | | Assumed output tokens | your `max_tokens` (or `max_completion_tokens`); 4096 when omitted, and written into the request; more than 32,768 is a 400 | | Images (`/v1/images/generations`) | assumed output = `n` × the model's per-image ceiling | | Rounding | reserve rounds **up**; settlement rounds **down** — both in your favor | The reserve exists only while the request is in flight: at settlement it is released and replaced by the observed cost, computed from provider-reported usage at the prices pinned when the request was admitted. Your balance and caps are only ever debited for observed usage — with one exception: an attempt whose outcome the gateway could not observe (`unknown_after_crash`) settles at the reserved maximum, because the call may have been billed upstream (see [Pricing & billing](/docs/pricing)). ## Balance exhaustion Every request must be coverable in the worst case: the org balance has to cover the worst-case reserve of *everything the org has in flight*, this request included. When it cannot, the refusal is the same `insufficient_quota` 429 with a different message: ```json { "error": { "code": "insufficient_quota", "type": "insufficient_quota", "message": "insufficient marketplace credits for authorized house fallback", "metadata": { "error_type": "insufficient_quota", "origin": "gateway_admission", "retryable": true } }, "request_id": "..." } ``` To tell the two `insufficient_quota` cases apart in code: `x-tm-cap-reset` is present **only** on cap breaches. > [!TIP] > Signup does not add free credit. Buy credits in [Billing](https://app.routerplus.com/console/billing); eligible purchases may receive matching promotional credit. If calls return `insufficient_quota`, check your balance and spend caps there. ## Handling the two 429s Both arrive as 429, so the OpenAI SDK raises `RateLimitError` for both — switch on `error_type`, because the correct reactions are opposite: ```python import os, time from openai import OpenAI, RateLimitError client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) def create(**kw): try: return client.chat.completions.create(**kw) except RateLimitError as e: kind = e.response.json()["error"]["metadata"]["error_type"] if kind == "rate_limit": time.sleep(int(e.response.headers.get("retry-after", "10"))) return client.chat.completions.create(**kw) # one retry raise # insufficient_quota: retrying cannot help — credits or a raised cap ``` ```typescript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY }); try { await client.chat.completions.create({ model: "claude-haiku-4-5", messages: [{ role: "user", content: "hello" }], }); } catch (err) { if (err instanceof OpenAI.APIError && err.status === 429) { const kind = (err.error as { metadata?: { error_type?: string } })?.metadata?.error_type; if (kind === "rate_limit") { // wait err.headers["retry-after"] seconds, retry once } else { // insufficient_quota — stop; x-tm-cap-reset (cap breaches only) says when a cap reopens const capReset = err.headers?.["x-tm-cap-reset"]; } } } ``` ## Other edges Being straight about the edges, so you can plan around them rather than discover them: - **No daily or weekly caps.** The only spend window is the UTC calendar month. If you need a tighter blast radius today, a low per-key cap on a purpose-made key is the tool. - **No self-serve RPM raises.** Trial keys are 20 RPM, console keys 300 (20 until the organization's first purchase); a limit you set can only lower a key's RPM. - **Sign-ups that look automated or abusive are refused.** The page says that unusual activity was detected; contact support if this happens to you. Existing accounts sign in without limit. - **Request size**: bodies over 10 MB get a 413 `request_too_large`. - **Image response size**: an upstream image response over 32 MiB is a 502 `gateway_error` and bills nothing — lower `n` or the quality. It is not a 413. - **An image request holds one admission slot for its whole render** (typically 10–60 s), because there is no first byte to commit on. When any of this changes it will change here first — this page is kept in lockstep with the enforcing code. --- Source: https://app.routerplus.com/docs/data-policy.md # Data policy The marketplace sits between your application and model providers. Its data posture is simple: **content-free by design**. Your prompts and the model's responses pass through the gateway on their way to and from the provider — they are never written down. What the marketplace records about a request is token *counts* and metadata: ids, timings, prices, error codes. This is not a configuration option or a retention window. There is no code path in the gateway that persists request or response content, so there is nothing to opt out of and nothing to delete later. One product is different by design, and says so: [Optimize](/docs/model-search) stores the eval cases and candidate answers of a search, for your organization only, until you delete the search. Nothing else holds content. ## The content-free rule Every store on the request path holds metadata only: | Store | What it holds | Content? | |---|---|---| | Process logs | request id, surface, model, outcome, HTTP status — every buyer-controlled string sanitized first | Never | | Telemetry (analytics) | ids, org/key ids, provider, deployment, model, region, attempt ordinal, outcome, error code, statuses, timestamps, latency, five token counts, request/response *byte sizes*, cost, provider request id (house routes only; see below) | Never | | Billing ledger | per-attempt row: the prices in force, reserved maximum, settled cost, token counts, HTTP status, error code, provider request id (house routes only; see below), timestamps, route and admission context | Never | | Rendered pages | model directory, model pages, console, status — all render from metadata | Never | Some details worth knowing: - **Model names appear in logs** (they are routing metadata), but every buyer-controlled string is sanitized before it touches a log line: control characters are stripped and length is capped, so request fields cannot forge log entries. - **Error bodies never echo your input.** Every error body is the gateway's own typed shape; a provider's raw error body is never relayed. When a refusal is translated across dialects or surfaced mid-stream, the gateway emits only the typed error class. - **Byte sizes are recorded, bodies are not.** Telemetry keeps `request_bytes` and `response_bytes` so throughput is measurable without storing a single byte of what was said. - **Image bytes and revised prompts are content.** On [`/v1/images/generations`](/docs/api-images) the prompt, the base64 image and any `revised_prompt` are never logged, never journaled and never stored — only the response byte size is. - **Videos are content too.** On [`/v1/videos`](/docs/api-videos) the prompt and the video file are never logged, journaled or stored. The gateway keeps the job's id, model, length, shape, status and charge, so you can read the job and fetch the file later; the file itself streams through from the provider when you download it. A provider's error text is reduced to a code before anything keeps it. - **The playground keeps your conversations, images and videos in your browser**: conversations in local storage, images and videos in IndexedDB, filed by your sign-in, the conversation and the reply. They are never uploaded. Deleting a chat, or **Clear all**, deletes them. - **A playground decision is content-free too.** The state and the answers pass through the control plane and the gateway to the decision model's provider, and are never logged or stored. The gateway keeps what it keeps for every request: model, provider, token counts and cost. - **Raw API keys never appear anywhere** — not in logs, not in the ledger, not in error bodies. The gateway authenticates by comparing a SHA-256 hash. - **The provider request id is kept on house routes only.** It is the reference number a provider gives one call, so a failed call can be traced with that provider. A house route runs on the marketplace's own provider account. A call through a provider credential you connect (BYOK) never keeps the id. A provider can put the API key into that header, so the gateway drops an id that contains the key it sent, or any 16 characters of the key in a row. It also drops an id longer than 256 characters. > [!NOTE] > One debug exception exists, and it is disclosed rather than hidden: setting > `TM_UNSAFE_LOG_BODIES=1` prints request bodies to stdout for local debugging. > The flag hard-refuses to activate outside `local`/`demo` regions — a > production deployment that sets it gets a logged warning and no body logging. > It never touches telemetry or the ledger in any region, and an image prompt > never prints even under it. The rule is enforced by code, review and tests: the integration suite plants sentinel content in playground chats, image and video prompts, and asserts it never appears in the control plane's or the gateway's logs, or in Postgres. ## What we do store Running an account requires a small amount of real data: | Data | Why | Form | |---|---|---| | Email address | Sign-in and billing receipts | Stored canonicalized (lowercased; gmail dots and `+tags` collapsed, `+tags` stripped elsewhere) so aliases share one account and promotional-credit eligibility | | Members' emails | Organization membership and roles | Canonicalized the same way | | API keys | Gateway authentication | SHA-256 hash plus the first 12 characters (so the console can show you *which* key). The raw key is shown exactly once, at creation | | Legacy and operator-issued email-link tokens | Redeeming previously issued sign-in links | SHA-256 hash only, single use, 30-minute expiry | | Clerk identity and session references | Connect verified sign-ins to the same marketplace account and support sign-out | Clerk issuer, user/session ids, verified canonical email, organization id and session timestamps | | Authentication email delivery IDs | Prevent duplicate sends when the Clerk email relay is enabled | Message id, outcome and timestamp only; no email body, sign-in link or recipient in this table | | Per-attempt metadata | Billing and the usage you read back | The ledger row described above | | Provider credentials you connect (BYOK) | Routing through your own OpenAI, Anthropic, Azure or Bedrock account | Envelope-encrypted; never returned by any API or page. A Bedrock connection stores a role ARN, never AWS keys | | Braintrust API key (Optimize) | Reading your evals | Encrypted; only its last four characters are ever shown | | Optimize searches | Comparing models on your evals | The eval cases and the candidates' answers, scoped to your organization, deleted with the search | | Video jobs | Following a render across reloads | Job id, model, length, shape, status and charge — never the prompt or the file | | App attribution | Optional `HTTP-Referer` / `X-Title` headers you send | Stored as-is in telemetry — send them only if you want your app identified | The marketplace does not store passwords. Clerk handles passwords, email verification and Google sign-in. Google sign-in requests basic identity information (OpenID, email and profile); it does not request access to Gmail, Drive or Calendar. The marketplace uses the verified email to find your existing account, including accounts originally created through an email link. Clerk processes authentication data under its [privacy policy](https://clerk.com/legal/privacy). Authentication emails can be delivered through Resend, including Clerk-generated email links. The relay verifies the sender's signature and does not log or persist the email body or sign-in link. Resend processes delivery data under its [privacy policy](https://resend.com/legal/privacy-policy). ## Account authentication Public signup and sign-in use Clerk in the browser at [Sign up](https://app.routerplus.com/signup) and [Sign in](https://app.routerplus.com/login). Complete the email-verification and authentication steps shown there. The local magic-link submission and programmatic signup routes have been removed. Existing accounts and API keys are preserved; Clerk connects a verified email to the same marketplace account. After Clerk verifies a sign-in, the marketplace issues its own seven-day signed session cookie (`HttpOnly`). Marketplace **Sign out** revokes that cookie and the linked Clerk session. Clerk-backed sessions are rechecked during authenticated requests, using cached proof for at most 30 seconds. A revoked or expired Clerk session can no longer authorize new marketplace requests after that cache expires. Other devices stay signed in. Beside the session cookie the browser holds a sign-in hint (`tm_in`, the same expiry, readable by script) that says only "signed in", and a display cache (`localStorage["tm-acct"]`: your email and your balance as the last page showed them). They let every page draw your account menu and credits at once. Neither opens anything, and both are cleared on sign out. Previously issued marketplace email links and operator-issued links use stored SHA-256 token hashes; a leaked database row is not a working link. These links are single-use and expire. They do not provide a public account-creation route. See [Authentication](/docs/authentication) for browser onboarding and API keys. ## What providers do with your prompts The marketplace not storing content does not mean the *provider* serving your request stores nothing. Every provider in the catalog must declare its prompt retention in its manifest before it can serve traffic: ```json { "provider": { "id": "anthropic", "privacy_policy_url": "https://www.anthropic.com/legal/privacy", "prompt_logging": "retained" } } ``` `prompt_logging` is a closed enum — `"none"` or `"retained"` — and a manifest without it (or without a privacy policy URL) fails conformance and never goes live. The flag is surfaced where you make decisions: - On every **model page**: the Data policy card says the marketplace's logs are content-free and how many of the model's providers offer zero data retention, and the providers table has a **Retention** column per provider. - In the machine-readable catalog, per provider of each model, and per host behind an aggregator as a `zdr` flag: ```bash curl -s https://app.routerplus.com/api/models.json | python3 -c \ 'import json,sys; [print(m["id"], "—", [(p["id"], p["prompt_logging"]) for p in m["providers"]]) for m in json.load(sys.stdin)["models"]]' ``` Every provider in the catalog today declares `"retained"`, per their published policies, except Fastino, Cloudflare Workers AI, Perplexity and RouterPlus. Levanto, which serves the decision model Sage, keeps no ordinary request, but keeps the full request of a failed or slow one for 30 days to debug it, so it declares `"retained"` too. Fastino, which serves the decision model GLiNER-2.5-Decide, declares `"none"`. Fastino keeps an inference unless the request says `store: false`, and the gateway sends `store: false` on every request. Fastino's terms let it train on inputs unless a team has Zero Data Retention on; our Fastino team has it on. So Fastino keeps nothing and trains on nothing. Cloudflare Workers AI, which serves the decision models Clef and Clef-flash, declares `"none"`, on Cloudflare's own statements. Of Clef, Cloudflare says: "we don't read, store, or train on your requests or responses (unless you want to use our fine-tuning product)". We do not use that product. Cloudflare's Workers AI data-usage page says that Cloudflare does not use a customer's content to train any AI model on Workers AI, or to improve any Cloudflare or third-party service, without the customer's explicit consent. It also says that Cloudflare stores the content only if the customer uses a Cloudflare storage service (R2 or KV, for example) with Workers AI. We use none. The gateway calls the Workers AI REST API directly. It never sends Cloudflare's AI Gateway header (`cf-aig-gateway-id`), so our requests do not go through AI Gateway, whose logs are on by default and can hold the full prompt and response. Cloudflare's replies to our calls carry none of AI Gateway's `cf-aig-*` headers. Perplexity, which serves the decision model Perplexity Decider v1 27B, declares `"none"`, on Perplexity's own statements. Its API FAQ says: "We do not retain any query data sent through the API and do not train on any of your data." It also says that the API has zero day retention of user prompt data by default, never used for AI training, and that its compute is hosted on Amazon Web Services in North America. Perplexity keeps the usage metadata it bills from, such as the number of requests and tokens (its API FAQ). The gateway calls Perplexity's API directly, with our one Perplexity key. Bespoke Labs, which serves the decision model Bespoke Nimble v3, declares `"retained"`: Bespoke keeps API inputs, outputs and request logs for up to 30 days, to debug the service and check for abuse, and then deletes them. It also keeps the metadata it bills from (token counts, times, key id, status). Bespoke does not train on API content unless an organization opts in, and ours does not. OpenRouter publishes no retention terms for the free endpoint that serves the decision model Mercury Decide, so that deployment (`openrouter-decisions-free`) declares `"retained"`, like `openrouter-decisions`. For a model served through OpenRouter, `served_by` lists the hosts that actually run it and which of them OpenRouter classifies as zero data retention; a request can pin a host with `provider.upstream` (see [Routing policies](/docs/routing-policies)). We report the flags; we do not editorialize them. ### RouterPlus RouterPlus is our own decision-model service. It serves our models Decider 2B and Kev 4B, and this page is its data policy. Its GPUs run in the United States. RouterPlus writes no request logs. It keeps a cache of its answers: each answer and its usage, for 600 s (10 minutes). The cache key is the RouterPlus API key, the model and the exact request body, byte for byte, so only an identical request gets a cached answer. A cache entry is never shared across RouterPlus keys, but the gateway sends every request with our one RouterPlus key: an identical request from any of our buyers within 600 s can get the cached answer, which repeats your option names. The cache is all it keeps, so the listing declares `"none"`. RouterPlus does not use your states, questions or answers for training. The marketplace side is the same as for every provider: content-free, as this page describes. ## Supply policy: where your requests can land Models from PRC-domiciled labs (DeepSeek, Moonshot, Z.ai, MiniMax, Tencent, Xiaomi) are listed through OpenRouter, not through those labs' own APIs: the catalog has no first-party deployment for them. Which host runs a given request is OpenRouter's choice among the hosts `served_by` lists for the model, unless you pin one with `provider.upstream`. This gets revisited only if a buyer needs first-party pricing or features the hosts lack — and it would be revisited openly, not quietly. ## Reading your own metadata back Everything the marketplace knows about your requests, you can read back with your API key: ```bash curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usage curl -s -H "Authorization: Bearer $TM_API_KEY" "https://api.routerplus.com/v1/generation?id=" ``` Both endpoints read the billing ledger — balance, attempts, token counts, settled cost. Neither can return message content, because none was stored. > [!TIP] > If you need to audit a specific request, keep the `x-request-id` response > header from the original call — it is the lookup key for `/v1/generation`, > and the id the marketplace will use if you write in about a billing dispute. See [Authentication](/docs/authentication) for account and API-key access, and [Errors](/docs/errors) for the error taxonomy referenced above. --- Source: https://app.routerplus.com/docs/api-chat-completions.md # POST /v1/chat/completions The OpenAI-compatible surface. Point any OpenAI SDK (or OpenAI-compatible client) at `https://api.routerplus.com/v1` with your marketplace key and it works — including against Claude models. Every chat model in the catalog is callable from this endpoint no matter which provider serves it; image models answer on [POST /v1/images/generations](/docs/api-images) only, and video models on [POST /v1/videos](/docs/api-videos) only; the gateway translates requests, streams, and errors between dialects. The Anthropic-native twin of this endpoint is [POST /v1/messages](/docs/api-messages). ``` POST https://api.routerplus.com/v1/chat/completions ``` ## Authentication | Header | Format | |---|---| | `Authorization` | `Bearer tm_vk_...` (the standard OpenAI-style header) | | `x-api-key` | `tm_vk_...` — also accepted, same key | A missing or invalid key returns 401 with `error_type: "auth"`. Optionally send `HTTP-Referer` and `X-Title` to identify your app (content-free attribution). ## Request JSON body, capped at 10 MB (413 `request_too_large` above that). | Parameter | Type | Behavior | |---|---|---| | `model` | string, required | A catalog id — see [Model discovery](#model-discovery). Unknown ids are an honest 404, never a silent substitute. A deployed endpoint from [Optimize](/docs/model-search) (`tm/-v`) is served here too, with a key of the organization that owns it, and so is a [dedicated endpoint](/docs/dedicated-endpoints) (`/`). | | `messages` | array, required | Roles `system`, `developer`, `user`, `assistant`, `tool`. Content: a string, or `text` parts, plus `image_url`, `file`, `input_audio` and `video_url` parts for a model that takes them (`architecture.input_modalities` in [`GET /v1/models`](/docs/api-models)). A part the model does not take is a typed 400 naming the part. Media parts reach the provider only when an OpenAI-dialect provider serves the request; on a route that needs translation they are a typed 400 naming the part. | | `stream` | boolean | Server-sent events; see [Streamed response](#streamed-response). | | `max_tokens` | number | Output cap, 1 to 32,768. **When you send neither `max_tokens` nor `max_completion_tokens`, the gateway sets `max_tokens: 4096`** on the request it forwards, whichever provider serves it. A value above 32,768 is a 400. | | `max_completion_tokens` | number | Honored as an alias. Sending both with different values is a 400. | | `temperature` | number | Passed through. Above 1 it cannot be translated to an Anthropic-dialect provider — see [Wire compatibility](/docs/compat). | | `top_p` | number | Passed through. | | `stop` | string \| string[] | Becomes Anthropic `stop_sequences` on cross-dialect routes. | | `tools` | array | `type: "function"` tools only. A tool without `parameters` gets an empty object schema on Anthropic-dialect routes. | | `tool_choice` | string \| object | `"auto"`, `"none"`, `"required"`, or `{"type":"function","function":{"name":"..."}}`. | | `parallel_tool_calls` | boolean | `false` becomes `disable_parallel_tool_use` on Anthropic-dialect routes. | | `n` | number | **Rejected when > 1** — 400 `invalid_request`. See below. | | `stream_options` | object | Usage is always on. An explicit `{"include_usage": false}` is honored: you are billed the same, but no usage chunk is written to you. Otherwise the object reaches an OpenAI-dialect provider with `include_usage` set, and is dropped and recorded on a translated route. | | `provider` | object | Routing controls read by the gateway and never forwarded: `require_parameters`, `order`, `only`, `ignore`, `allow_fallbacks`, `upstream` and the rest — see [Routing policies](/docs/routing-policies). An unknown control is a 400. | | `session_id` | string | A session id: the gateway sends each provider its own session field (`prompt_cache_key` to OpenAI and Azure, `session_id` to OpenRouter). The headers `x-session-id`, `x-session-affinity` and `x-claude-code-session-id` work too. See [Sessions and prompt caching](/docs/compat#sessions-and-prompt-caching). | | `prompt_cache_key` | string | OpenAI's cache key: sent unchanged to OpenAI and Azure, and to OpenRouter. | | `prompt_cache_retention` | string | `"in_memory"` or `"24h"`: sent unchanged to OpenAI and Azure; dropped and recorded elsewhere. | | `safety_identifier` | string | Your end-user id. On a house route the gateway sends a hash of it with your organization; see [Sessions and prompt caching](/docs/compat#sessions-and-prompt-caching). | | `cache_control` | object | Automatic prompt caching, as on the Anthropic API: sent to Anthropic and OpenRouter; dropped and recorded on OpenAI and Azure. A field that Anthropic would refuse is dropped and recorded on every route (see [Wire compatibility](/docs/compat)). A `cache_control` mark on a `text` part is kept for Anthropic and OpenRouter, and passes unchanged to OpenAI and Azure, which ignore it. | A top-level `api_key`, `base_url`, `connection_id` or similar credential or destination field is a 400: keys and endpoints are configured in the console, never sent in a request. ### Everything else The table above is the **portable** set — it survives translation to any provider. What happens to parameters outside it depends on who serves the request: when the serving provider speaks the OpenAI dialect, the OpenAI chat parameters (`response_format`, `reasoning_effort`, `seed`, `user`, `logprobs`, `top_logprobs`, `frequency_penalty`, `presence_penalty`, `logit_bias`, `metadata`, `store`, `prediction`) reach it as you sent them; when translation to the Anthropic dialect is needed, `response_format`, `reasoning_effort`, `prediction`, `audio` and `modalities` are a typed 400 (they change what the model produces), and other unknown parameters are dropped — never silently mutated into something else. Any top-level key outside those sets is dropped and recorded on every route. Every drop is named in the `x-tm-dropped-params` header; `"provider": {"require_parameters": true}` turns a drop into a 400 instead. The full tables are on [Wire compatibility](/docs/compat). Unsupported **content** is different from unsupported parameters: it is never dropped, because dropped content would still be billed upstream. A part the model does not take — an image to a model that takes text only, say — is a typed 400 naming the exact field (e.g. `messages[0].content[1]`), on every route, before anything is billed or sent. On routes that need translation to the Anthropic dialect, image parts, audio parts, and unknown content-part types are a typed 400 naming the field too. When the serving provider speaks the OpenAI dialect and the model takes the part, content passes through as sent. ### Why `n > 1` is rejected The gateway's normalizer emits exactly one choice. Accepting `n: 3` and billing you for a garbled single-choice response would be dishonest, so the request is refused up front: ```json { "error": { "code": "invalid_request", "message": "n>1 is not supported in v1; request a single choice", "type": "invalid_request_error", "metadata": { "error_type": "invalid_request" } }, "request_id": "..." } ``` ## Non-streamed response A standard `chat.completion` object. On translated responses the `id` is `chatcmpl-`; when an OpenAI-dialect provider serves the request, the body passes through with the provider's own `id`. For the `/v1/generation?id=` audit, use the `x-request-id` response header — it carries the bare request id on every route. On a [closed dedicated endpoint](/docs/dedicated-endpoints#closed-endpoints) the body keeps a fixed set of fields on every route: `model` is the endpoint id, the `id` is `chatcmpl-` and the request id without dashes, and provider fields such as `system_fingerprint` and `native_finish_reason` are dropped. ```json { "id": "chatcmpl-1c9a7b2e-...", "object": "chat.completion", "created": 1757000000, "model": "claude-sonnet-4-5", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Hello!" }, "finish_reason": "stop", "native_finish_reason": "end_turn" } ], "usage": { "prompt_tokens": 12, "completion_tokens": 5, "total_tokens": 17, "prompt_tokens_details": { "cached_tokens": 0, "cache_write_tokens": 0 }, "cost": 0.000111 } } ``` Usage fields (the billing contract — see [Wire compatibility](/docs/compat)): | Field | Meaning | |---|---| | `prompt_tokens` | All billed input tokens — **cache-inclusive**: uncached input + cache reads + cache writes. | | `prompt_tokens_details.cached_tokens` | Cache reads (billed at the cached rate). | | `prompt_tokens_details.cache_write_tokens` | Cache writes (billed at their own rate). Always present on billed responses, 0 when none. | | `completion_tokens` | Output tokens. Reasoning tokens are a subset, reported in `completion_tokens_details.reasoning_tokens` when the provider reports them (streams always carry the field). | | `cost` | Settled cost in USD — billed keys only. Computed with the exact integer micro-USD math the ledger settles with, covering the **full request** debit including attempts that failed over before your answer started. You can recompute your bill from the wire. | `finish_reason` is one of `stop`, `length`, `tool_calls`, `content_filter`. On responses translated from an Anthropic-dialect provider, `choices[0].native_finish_reason` preserves the provider-raw stop reason (e.g. `end_turn`, `stop_sequence`) — a documented extension. A model's thinking, when the provider returns it, is in `message.reasoning_content`. ## Streamed response With `stream: true`, SSE frames in `chat.completion.chunk` shape, in a fixed order: 1. A role-priming delta (`{"role":"assistant","content":""}`) as soon as the upstream proves alive. 2. Content deltas: `delta.content`, `delta.reasoning_content` (thinking models), and `delta.tool_calls` fragments addressed by `index` (id and `function.name` arrive on each call's first fragment). 3. A finish chunk with `finish_reason` set, plus `native_finish_reason` on translated streams. 4. A usage chunk (empty `choices`) — always sent, no `stream_options` needed: ```text data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":1757000000,"model":"claude-haiku-4-5","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":5,"total_tokens":17,"prompt_tokens_details":{"cached_tokens":0,"cache_write_tokens":0},"completion_tokens_details":{"reasoning_tokens":0},"cost":0.000037}} data: [DONE] ``` 5. `data: [DONE]`. When an OpenAI-dialect provider serves the request, its stream is relayed record for record, with `usage.cost` added to its usage chunk; the shapes above are what a translated stream looks like. On a closed dedicated endpoint each chunk keeps only the fixed set of fields, and a provider's own SSE comment lines are dropped. Full detail on [Streaming](/docs/streaming). During post-start silences of 15 s or more (slow reasoning models), the gateway emits an SSE comment (`: processing`) so proxies don't kill an idle-but-healthy stream — spec-ignorable, and your SDK already ignores it. > [!WARNING] > If the serving provider dies **after** your answer started, the stream carries one terminal error chunk — a full chunk envelope with `finish_reason: "error"` and a top-level `error` object carrying `metadata.error_type` — and nothing after it. **No `[DONE]` follows an error.** The answer is never silently restarted on a different provider mid-stream; failover happens only before the first byte of output, invisibly. The same chunk ends a stream that reaches the request deadline (450 s from dispatch, or a routing policy's `timeout_ms`). ## Errors Every error uses the OpenAI envelope with a stable `error.metadata.error_type` (`auth`, `rate_limit`, `insufficient_quota`, `model_unavailable`, `invalid_request`, `context_overflow`, `content_policy`, `upstream_error`, `gateway_error`, ...). A provider's own error body is never relayed; when a provider rejected the request, the class says so and the `x-tm-upstream-status` header carries its status. Full table with remediations: [Errors](/docs/errors). Out-of-credits is the OpenAI-standard **429 `insufficient_quota`** (your SDK already recognizes it), covering an empty balance or a monthly spend cap — cap breaches include the exact UTC reset time in the message and the `x-tm-cap-reset` header. First trial keys shown after verified sign-in have a 20 RPM limit, and until an organization buys credit its keys together run at 20 RPM; 429 `rate_limit` carries `retry-after`. > [!TIP] > Billed requests reserve worst-case cost before dispatch: the request's size in bytes as the input estimate (one token per byte, plus 65,536 per image part; an image sent inline as base64 counts the 65,536 only, not its bytes) and `max_tokens` (4096 when omitted) as the output, at the model's rates, rounded up. If your balance can't cover it — including your org's other in-flight reservations — you get `insufficient_quota` before any provider is called. Settlement is on observed usage only, rounded down, so a generous `max_tokens` never costs extra — but it can make a low balance refuse early. Set it realistically. ## Response headers | Header | Meaning | |---|---| | `x-request-id` | Gateway request id; joins `/v1/usage` and `/v1/generation?id=`. | | `x-tm-provider` | Deployment that served the request; `routerplus` on a closed dedicated endpoint. | | `x-tm-attempts` | Physical dispatches, failovers included. | | `x-tm-upstream-status` | The provider's own HTTP status. | | `x-tm-error-code` | Canonical error class, on failures. `x-tm-error-origin` says where the failure came from, and `x-tm-limit-scope` / `x-tm-limit-kind` name a limit that refused the request — see [Errors](/docs/errors). | | `x-tm-upstream-model` | Present when the serving provider knows this model under its own id (aggregators): the id the gateway actually sent. The response's `model` field echoes the provider's id. Absent on a closed dedicated endpoint. | | `x-tm-served-by` | Present when the serving route is an aggregator that names the provider it used (for example `Amazon Bedrock`). Absent on a direct route: there `x-tm-provider` is the provider. Absent on a closed dedicated endpoint. | | `x-tm-dedicated-endpoint`, `x-tm-route-role` | On a [dedicated endpoint](/docs/dedicated-endpoints): its id, and `primary` or `fallback`. | | `x-tm-queue-ms` | On a dedicated endpoint with a [paced limit](/docs/dedicated-endpoints#the-paced-limit): how long the request waited for its turn, in milliseconds. | | `x-tm-dropped-params` | Comma-joined field paths of any parameters the gateway stripped for this route (D8 §2: drops are recorded, never silent). Absent when nothing was dropped. Send `"provider": {"require_parameters": true}` to get a typed 400 instead of any drop. | | `x-tm-route-plan-id` | On billed requests: the id of the route plan that chose the deployments. | | `x-tm-admission-mode`, `x-tm-remaining-rpm`, `x-tm-remaining-tpm` | On billed requests: how the request was admitted and the smallest allowance left across the limits it claimed — see [Limits and capacity](/docs/admission). | ## Examples ```bash curl -N https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"claude-haiku-4-5","stream":true,"max_tokens":100,"messages":[{"role":"user","content":"Say hello."}]}' ``` Python — the official `openai` SDK with the base URL swapped: ```python import os from openai import OpenAI client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) resp = client.chat.completions.create( model="claude-sonnet-4-5", # a Claude model over the OpenAI wire — the gateway translates max_tokens=200, messages=[{"role": "user", "content": "Say hello in one sentence."}], ) print(resp.choices[0].message.content) print(resp.usage) # prompt/completion tokens, cache splits, cost (USD) ``` TypeScript — the official `openai` package: ```typescript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY, }); const stream = await client.chat.completions.create({ model: "gpt-4o-mini", stream: true, messages: [{ role: "user", content: "Count to five." }], }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); if (chunk.usage) console.log("\n", chunk.usage); // final chunk: tokens, cache splits, cost } ``` ## Model discovery `GET https://api.routerplus.com/v1/models` (authenticated) returns the catalog in the OpenAI list shape — one row per model id — and your organization's dedicated endpoints. The public directory with prices and context lengths is `https://app.routerplus.com/api/models.json`. A model id that is not listed returns 404 `model_unavailable` naming the id you asked for; if every deployment serving a listed model is cooling down after failures, you get a 503 `gateway_error` with `retry-after: 5` and `x-tm-limit-kind: health` instead of a fake 404 — retry, don't fix. See also: [POST /v1/messages](/docs/api-messages) · [Errors](/docs/errors) · [Wire compatibility & billing contract](/docs/compat) · [Quickstart](/docs/quickstart) --- Source: https://app.routerplus.com/docs/api-messages.md # POST /v1/messages The Anthropic-compatible surface. Point the Anthropic SDK (or Claude Code) at `https://api.routerplus.com` with your marketplace key and it works — including against GPT models. Every chat model in the catalog is callable from this endpoint regardless of which provider serves it; image models answer on [POST /v1/images/generations](/docs/api-images) only, and video models on [POST /v1/videos](/docs/api-videos) only; the gateway translates requests, streams, and errors between dialects. The OpenAI-native twin is [POST /v1/chat/completions](/docs/api-chat-completions). ``` POST https://api.routerplus.com/v1/messages ``` Query strings are tolerated — clients that append them (Claude Code sends `/v1/messages?beta=true`) route normally. ## Headers | Header | Required | Behavior | |---|---|---| | `x-api-key` | one of these two | Your marketplace key, `tm_vk_...` — the Anthropic-native header. | | `Authorization` | | `Bearer tm_vk_...` also accepted, same key. | | `content-type` | yes | `application/json`. | | `anthropic-version` | no | Accepted for SDK compatibility; the gateway speaks `anthropic-version: 2023-06-01` to Anthropic upstreams regardless of what you send. | | `anthropic-beta` | no | Forwarded when an Anthropic-dialect provider serves the request, for the values that do not change how a token is billed: `prompt-caching`, `token-efficient-tools`, `fine-grained-tool-streaming`, `interleaved-thinking`, `claude-code`, `oauth` and `computer-use` prefixes. Any other value is dropped and recorded in `x-tm-dropped-params` as `header:anthropic-beta:`. No effect on other routes. | A missing or invalid key returns 401 in the Anthropic error shape with `error_type: "auth"`. Optionally send `HTTP-Referer` and `X-Title` for content-free app attribution. ## Request JSON body, capped at 10 MB (413 `request_too_large` above that). | Parameter | Type | Behavior | |---|---|---| | `model` | string, required | A catalog id — any chat model, not just Claude. Unknown ids are an honest 404; an image model is a 400 naming `/v1/images/generations`, a video model a 400 naming `/v1/videos`. A deployed endpoint id (`tm/-v`) is served on `POST /v1/chat/completions` only and is a 400 here. | | `max_tokens` | number | Output cap, 1 to 32,768. **When you omit it, the gateway sets `max_tokens: 4096`** on the request it forwards, whichever provider serves it. A value above 32,768 is a 400. | | `messages` | array, required | Roles `user` and `assistant`. Content: a string, or blocks — `text`, `tool_use` (assistant), `tool_result` (user), `thinking` / `redacted_thinking` (assistant). `image` and `document` blocks are for a model that takes them (`architecture.input_modalities` in [`GET /v1/models`](/docs/api-models)): a block the model does not take is a typed 400 naming it, also inside a `tool_result`. `image` blocks are a typed 400 when translation to an OpenAI-dialect provider is needed; on Anthropic-dialect routes the messages pass through as sent. | | `system` | string \| array | A string or text blocks. | | `stream` | boolean | SSE streaming; see below. | | `temperature`, `top_p` | number | Passed through. | | `stop_sequences` | string[] | Becomes OpenAI `stop` on cross-dialect routes. | | `tools` | array | `{name, description?, input_schema}`. Anthropic server tools (computer use, web search, bash, text editor) pass through on Anthropic-dialect routes and are a typed 400 when an OpenAI-dialect provider would serve the request. | | `tool_choice` | object | `{"type":"auto"}`, `{"type":"any"}`, `{"type":"none"}`, or `{"type":"tool","name":"..."}`. `disable_parallel_tool_use: true` becomes `parallel_tool_calls: false` on OpenAI-dialect routes. | | `thinking`, `output_config` | object | Pass through on Anthropic-dialect routes; a typed 400 when translation to an OpenAI-dialect provider is needed — the gateway does not map them yet and will not drop them silently. | | `top_k`, `metadata` | | Pass through on Anthropic-dialect routes; dropped and recorded on OpenAI-dialect routes. On a house route, `metadata.user_id` becomes a hash with your organization; see [Sessions and prompt caching](/docs/compat#sessions-and-prompt-caching). | | `cache_control` | object | Automatic prompt caching, on the request or on blocks. Passes to Anthropic and OpenRouter (on Bedrock the request-level field becomes a mark on the last block); dropped and recorded on OpenAI and Azure. A request-level field that Anthropic would refuse is dropped and recorded on every route (see [Wire compatibility](/docs/compat)). On a house route, `ttl: "1h"` is removed (5 minutes) and recorded. | | `session_id` | string | A session id, also accepted as the header `x-session-id`, `x-session-affinity` or `x-claude-code-session-id`. It reaches OpenRouter as `session_id` and OpenAI or Azure as `prompt_cache_key`; Anthropic has no session input. See [Sessions and prompt caching](/docs/compat#sessions-and-prompt-caching). | | `mcp_servers`, `container` | | A typed 400 on every route: stripping them would change who runs what. | | `provider` | object | Routing controls read by the gateway and never forwarded: `require_parameters`, `order`, `only`, `ignore`, `allow_fallbacks`, `upstream` and the rest — see [Routing policies](/docs/routing-policies). An unknown control is a 400. | Any other top-level key is dropped and recorded on every route (the `x-tm-dropped-params` header names it), never mutated. A top-level `api_key`, `base_url` or `connection_id` is a 400. Unsupported **content** is never dropped: a block the model does not take, or one a translation cannot carry, is a typed 400 naming the exact field, because dropped content would still be billed upstream. A typed 400 from translation is decided per deployment: when another deployment of the model speaks the Anthropic dialect, the request goes there instead — see [Wire compatibility](/docs/compat). > [!NOTE] > Transcripts with `thinking` blocks round-trip safely. When a model served over the OpenAI dialect streams reasoning, this surface encodes it as `thinking` blocks — and when you echo that transcript back and the request routes to an OpenAI-dialect provider, assistant `thinking`/`redacted_thinking` blocks are dropped rather than rejected. The route that produced your transcript will never 400 on it. ## Streamed response With `stream: true`, the real Anthropic Messages event sequence — the same normalized encoding whichever provider serves the request. On a translated stream the message `id` is `msg_`, joinable with the `x-request-id` header and `/v1/generation?id=`. ```text event: message_start data: {"type":"message_start","message":{"id":"msg_1c9a7b2e-...","type":"message","role":"assistant","model":"claude-sonnet-4-5","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":0,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":0}}} event: content_block_start data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}} event: content_block_delta data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}} event: content_block_stop data: {"type":"content_block_stop","index":0} event: message_delta data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"input_tokens":12,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":5,"cost":0.000111}} event: message_stop data: {"type":"message_stop"} ``` Block types you may see: `text` (`text_delta`), `thinking` (`thinking_delta`, plus `signature_delta` for multi-turn thinking with tools), `redacted_thinking`, and `tool_use` (`input_json_delta` fragments). Streamed `stop_reason` is one of `end_turn`, `max_tokens`, `tool_use`, `refusal`. Usage fields, per the [billing contract](/docs/compat): `input_tokens` **excludes** cache reads and writes; `cache_read_input_tokens` and `cache_creation_input_tokens` carry the cache splits; `usage.cost` (USD, billed keys only) in the final `message_delta` is the full request debit, failed-over attempts included, computed with the exact math the ledger settles with. > [!NOTE] > Treat the final `message_delta`'s usage as authoritative. On same-dialect routes (a Claude model on this surface) the stream is relayed record-for-record, so `message_start` carries the provider's real token counts and its own message id. On cross-dialect routes (an OpenAI-dialect provider behind this surface) the gateway re-encodes the stream and `message_start`'s counts are zeros — the final `message_delta` always carries the real numbers on every route. During post-start silences of 15 s or more the gateway emits native `ping` events (`{"type":"ping"}`) so idle-but-healthy streams survive proxies. If the serving provider dies **after** output started, you get one terminal `error` event inside the 200 stream — carrying the Anthropic-native `type` plus a stable `error_type` — and nothing after it; the answer is never silently restarted on another provider. The same event ends a stream that reaches the request deadline (450 s from dispatch, or a routing policy's `timeout_ms`). One documented limit: a billed stream that dies mid-answer has no legal Anthropic wire slot for usage outside `message_delta`, so recompute that request's cost from `/v1/generation?id=`. ## Non-streamed response ```json { "id": "msg_1c9a7b2e-...", "type": "message", "role": "assistant", "model": "gpt-4o-mini", "content": [{ "type": "text", "text": "Hello!" }], "stop_reason": "end_turn", "stop_sequence": null, "usage": { "input_tokens": 12, "output_tokens": 5, "cache_read_input_tokens": 0, "cache_creation_input_tokens": 0, "cost": 0.000004 } } ``` `stop_reason` on translated responses is `end_turn`, `max_tokens`, `tool_use`, or `refusal`; when an Anthropic-dialect provider serves the request the body passes through with its native values (including `stop_sequence`) plus the injected `cost`. ## Errors Anthropic error shape with the native `type` string your SDK switches on (`authentication_error`, `rate_limit_error`, `billing_error`, `not_found_error`, ...) plus a stable `error_type` carrying the canonical class: ```json { "type": "error", "error": { "type": "not_found_error", "message": "model \"claude-9\" is not in the catalog; GET /v1/models lists what this key can serve", "error_type": "model_unavailable" }, "request_id": "..." } ``` Out-of-credits and spend-cap breaches are 429 with native type `billing_error` and `error_type: "insufficient_quota"` (cap breaches include the exact UTC reset time and an `x-tm-cap-reset` header). A key whose signup email is not verified yet gets 403 with native type `permission_error` and `error_type: "email_not_verified"`. Until an organization has added credit, its keys and the organization as a whole run at 20 RPM — 429 `rate_limit` with `retry-after`. A provider's own error body is never relayed; when a provider rejected the request, the class says so and `x-tm-upstream-status` carries its status. Full table: [Errors](/docs/errors). Every response carries `x-request-id`; served requests add `x-tm-provider` and `x-tm-attempts`, failures add `x-tm-error-code` (and `x-tm-error-origin` when a limit, a provider or our infrastructure refused the request), and `x-tm-upstream-status` appears whenever an upstream responded. The other headers are the same as on [POST /v1/chat/completions](/docs/api-chat-completions#response-headers). ## Examples ```bash curl -N https://api.routerplus.com/v1/messages \ -H "x-api-key: $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"gpt-4o-mini","stream":true,"max_tokens":100,"messages":[{"role":"user","content":"Say hello."}]}' ``` Yes — a GPT model over the Anthropic wire. Python, with the official `anthropic` SDK and the base URL swapped (the SDK appends `/v1/messages` itself): ```python import os from anthropic import Anthropic client = Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"]) with client.messages.stream( model="claude-sonnet-4-5", max_tokens=300, messages=[{"role": "user", "content": "Count to five."}], ) as stream: for text in stream.text_stream: print(text, end="") print("\n", stream.get_final_message().usage) # tokens, cache splits, cost ``` TypeScript — `@anthropic-ai/sdk`: ```typescript import Anthropic from "@anthropic-ai/sdk"; const client = new Anthropic({ baseURL: "https://api.routerplus.com", apiKey: process.env.TM_API_KEY, }); const stream = client.messages.stream({ model: "claude-haiku-4-5", max_tokens: 200, messages: [{ role: "user", content: "Say hello." }], }); stream.on("text", (t) => process.stdout.write(t)); console.log("\n", (await stream.finalMessage()).usage); ``` Claude Code works unmodified: ```bash ANTHROPIC_BASE_URL=https://api.routerplus.com ANTHROPIC_AUTH_TOKEN=$TM_API_KEY claude ``` ## Model discovery `GET https://api.routerplus.com/v1/models` returns the catalog; send an `anthropic-version` header (the Anthropic SDK does) to get the Anthropic list shape instead of the OpenAI one. That shape lists chat models only. The public directory with prices and context lengths is `https://app.routerplus.com/api/models.json`. See also: [POST /v1/chat/completions](/docs/api-chat-completions) · [Errors](/docs/errors) · [Wire compatibility & billing contract](/docs/compat) · [Migrate in one prompt](/docs/migration) ## POST /v1/messages/count_tokens An unbilled passthrough for Anthropic-dialect deployments — the same request shape as `/v1/messages`, answered by the provider's own tokenizer: ```bash curl -s https://api.routerplus.com/v1/messages/count_tokens \ -H "x-api-key: $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model": "claude-sonnet-5", "messages": [{"role": "user", "content": "hello"}]}' ``` ```json { "input_tokens": 8 } ``` The answer is the count only. The call creates no ledger attempt and costs nothing, but it counts against your key's rate limit and the provider's pool like any request. An unknown model is a 404; if only OpenAI-dialect deployments (or a Bedrock connection) serve the model, the response is a typed 400: the gateway will not fabricate a token estimate. --- Source: https://app.routerplus.com/docs/api-images.md # POST /v1/images/generations Generate images from a text prompt. This is the OpenAI Images API surface: point `client.images.generate` at `https://api.routerplus.com/v1` and the call works unchanged. The path `https://api.routerplus.com/images/generations` is an alias for the same handler, for SDKs that omit the `/v1` prefix. To try it without writing code, pick an image model in the [playground](/docs/playground). The route is deliberately small: generations only (no edits, no variations), non-streaming, and the image always comes back as base64 in the body. Everything else is a typed error, never a silent degradation. The GPT Image models are served by OpenAI direct. Google's Nano Banana family, ByteDance's Seedream and xAI's Grok Imagine are served through OpenRouter today and billed at the cost it reports — see [Models billed at reported cost](#models-billed-at-reported-cost). ## Authentication | Header | Format | |---|---| | `Authorization` | `Bearer tm_vk_...` (the standard OpenAI-style header) | | `x-api-key` | `tm_vk_...` — also accepted, same key | A missing or invalid key returns 401 with `error_type: "auth"`. Optionally send `HTTP-Referer` and `X-Title` to identify your app (content-free attribution). ## Request JSON body, capped at 10 MB (413 `request_too_large` above that). Every rule below is checked **before** anything is reserved or dispatched, so a refused request has no ledger row, no `x-tm-attempts` header, and no charge. | Parameter | Type | Behavior | |---|---|---| | `model` | string, required | An image model id — see [Model discovery](#model-discovery). A dated pin (`-20260421`, `-2026-04-21`, `-latest`) resolves to its catalog family for routing and pricing, and your exact string goes upstream. A **chat** model here is a 400 naming the chat routes; a video model a 400 naming `/v1/videos`. An id that is not an image model anywhere is a 404. Omitted or empty is a 400 naming `model`. | | `prompt` | string, required | Non-empty, at most 32 000 characters. | | `n` | integer | 1 to 4 on the GPT Image models; the OpenRouter models take 1 (Seedream 5.0 Lite up to 4). Defaults to 1. Ours is stricter than the SDK's 10: four buffered images already approach the response cap. | | `size` | string | `auto`, `1024x1024`, `1536x1024` or `1024x1536` on the GPT Image models. The OpenRouter models take `aspect_ratio` instead, and drop `size`. | | `aspect_ratio` | string | On the OpenRouter models: `1:1`, `16:9`, `9:16`, … — each model's own list. | | `resolution` | string | On most OpenRouter models: `512`, `1K`, `2K` or `4K`, per model. | | `seed` | integer | On Seedream. | | `quality` | string | `auto`, `low`, `medium`, `high` — plus `xhigh` and `max` on `gpt-image-2.5-flare` and `gpt-image-2.5-sunburst`; `low` or `medium` on Grok Imagine. | | `background` | string | `auto`, `transparent`, `opaque` on the GPT Image models. | | `moderation` | string | `auto` or `low` on the GPT Image models. | | `output_format` | string | `png`, `jpeg` or `webp` on the GPT Image models. | | `output_compression` | integer | 0 to 100 on the GPT Image models. | | `user` | string | Forwarded verbatim. Anything but a string is a 400. | | `response_format` | string | `b64_json` only, and it is a no-op. `url` is a 400 — see [Always base64](#always-base64). | | `stream`, `partial_images` | — | Absent, `false` and `0` are no-ops. Anything else is a 400: this route does not stream. | | `provider` | object | Routing controls read by the gateway and never forwarded — `require_parameters` and the rest, see [Routing policies](/docs/routing-policies). | The accepted values for `size`, `aspect_ratio`, `resolution`, `quality`, `background`, `moderation`, `output_format`, `output_compression` and `seed` are **per model** and published in `/api/models.json` under `supported_parameters`. A value outside a model's declared set is a 400 whose message starts with the field name and lists what that model accepts. One of these parameters that a model does not declare at all — `size` on an OpenRouter model, `quality` on Nano Banana — is dropped and recorded, like any unknown key. An explicit `null` means "not provided" for every optional parameter above, exactly like omitting the key: the OpenAI SDKs serialise an unset optional as `null`, so `client.images.generate(model=…, prompt=…, n=None, size=None)` works. A `null` is never forwarded upstream and never recorded as a drop. `model` and `prompt` are required, so `null` there is still a 400. ### Everything else `style` and any key not in the table are **dropped and recorded**, never forwarded and never guessed at (D8 §2). The dropped names are comma-joined in the `x-tm-dropped-params` response header, stored on the attempt's ledger row, and visible at `GET /v1/generation?id=`. To refuse instead of dropping, send `"provider": {"require_parameters": true}` — a would-be drop then becomes a 400 before any money is reserved. `provider` itself is consumed by the gateway and never forwarded. ### Always base64 The gateway always returns the image bytes inline as `b64_json`. There is no hosted URL: storing an image would make this service stateful and would put buyer content on our disks, which the [data policy](/docs/data-policy) forbids. `response_format: "url"` is therefore a typed 400, not a quietly different answer. Decode the base64 yourself — the examples below do it in three languages. ## Response HTTP 200, `application/json`, with OpenAI's own `ImagesResponse` relayed untouched except for one added field: ```json { "created": 1790000000, "data": [{ "b64_json": "iVBORw0KGgoAAAANSUhEUg..." }], "background": "opaque", "output_format": "png", "quality": "medium", "size": "1024x1024", "usage": { "input_tokens": 12, "input_tokens_details": { "text_tokens": 12, "image_tokens": 0 }, "output_tokens": 1056, "total_tokens": 1068, "cost": 0.0423 } } ``` | Field | Meaning | |---|---| | `data[].b64_json` | The image, base64-encoded. One entry per image the provider rendered — normally `n`; the gateway relays what it received and bills what was reported. | | `data[].revised_prompt` | The provider's rewritten prompt, when it sends one. Relayed, never stored. | | `usage.input_tokens` | Prompt tokens, billed at the model's `prompt` rate. | | `usage.output_tokens` | Image output tokens, billed at the model's `completion` rate. | | `usage.input_tokens_details` | Informational only — the bill is computed from the two counts above. | | `usage.total_tokens` | `input_tokens + output_tokens`. | | `usage.cost` | USD, the **full request debit** including any attempt that failed over before this one. `0` on BYOK. Absent on static dev keys. | The chat surface's `prompt_tokens_details` is never injected here; an images response keeps its own shape. > [!NOTE] > If the provider returns a 200 without a usage object, or with `input_tokens` and > `output_tokens` both 0, the gateway bills `n` × the model's > per-image ceiling with provenance `estimated` — the same amount the reservation held — > and writes those numbers into `usage` so your copy of the bill still adds up. The > provenance is visible at [GET /v1/generation](/docs/api-usage). ## Billing Same two SKUs as chat, same integer micro-USD math ([Pricing & billing](/docs/pricing)): text in bills at `prompt`, image out bills at `completion`. - **Reserve**, before any provider is contacted: prompt bytes ÷ 4 at the `prompt` rate, plus `n` × the model's per-image ceiling at the `completion` rate, rounded up. The chat path's `max_tokens` default plays no part here. - **Settle**, on the provider-reported counts, rounded down. The hold is released in full. Worked numbers on `gpt-image-1` (ceiling 6240 tokens, $5 / $40 per million): | Case | Amount | |---|---| | Reserve, `n: 1` | 6240 × $40/M ≈ $0.2496, plus the prompt estimate | | Reserve, `n: 4` | 4 × 6240 × $40/M ≈ $0.9984, plus the prompt estimate | | Settle, one medium 1024×1024 image | 1056 × $40/M = $0.04224, plus the prompt | **All-or-nothing.** A generation either arrives whole or costs nothing. A failed attempt, a response over the 32 MiB cap, and a 200 that carries no decodable image all settle at zero. Once a 200 has been received the gateway never re-dispatches, so one request can never buy two renders. **A cancelled render is billed at what the provider reported.** If you hang up mid-render the upstream call is *not* aborted: the render finishes, the usage is read, and the attempt settles the provider's own numbers — never the reservation, never a pretend $0. Nothing is written to your dead socket. > [!WARNING] > Reservations are large next to a trial balance. `n: 4` at a high quality holds about a > dollar, and your balance must cover every reservation your org has in flight. A > `429 insufficient_quota` on `n: 4` is correct behavior, not a bug — lower `n`, or add > credits. ## Errors | Case | Status | `error_type` | Billed | |---|---|---|---| | A field outside the model's accepted values, `n` outside 1–4, `stream`, `partial_images`, `response_format: "url"`, a chat or video model on this route | 400 | `invalid_request` | no row | | The model is listed but has no price row | 400 | `model_not_priced` | no row | | Not an image model in the catalog | 404 | `model_unavailable` | no row | | Request body over 10 MB | 413 | `request_too_large` | no row | | Key RPM or a shared limit, balance too low, or a monthly spend cap | 429 | `rate_limit` / `insufficient_quota` | no row | | OpenAI moderation refusal (`moderation_blocked`) | the provider's, usually 400 | `content_policy` | $0; never rerouted, never a health strike | | Any other provider 4xx | the provider's | `upstream_error` | $0; no failover | | Provider 5xx, a failed connection, or the 450 s budget | 5xx relayed / 502 / 504 | `upstream_error` / `model_unavailable` | $0 per failed attempt; fails over like chat | | Response over the 32 MiB cap | 502 | `gateway_error` | $0; no failover | | A 200 that is unparseable or has no `b64_json` | 502 | `upstream_error` | $0; no failover | | The budget ran out while reading the image body | 504 | `upstream_error` | $0; no failover | | Every deployment serving the model is cooling down | 503 | `gateway_error` | nothing dispatched; `retry-after: 5` | | The ledger journal, the shared limit store, or the admission lease is unavailable | 503 | `gateway_error` | nothing charged | Bodies follow the OpenAI error envelope with the canonical class in `error.metadata.error_type`, and the class is on every failure in `x-tm-error-code`. A provider's own error body is never relayed; its status is in `x-tm-upstream-status`. The full taxonomy and remediation table is in [Errors](/docs/errors). The reverse refusal also holds: an image model sent to `POST /v1/chat/completions`, `POST /v1/messages` or `POST /v1/videos` is a 400 `invalid_request` naming `/v1/images/generations`. ## Limits | Limit | Value | |---|---| | Images per request | `n` ≤ 4 | | Prompt length | 32 000 characters | | Request body | 10 MB → 413 `request_too_large` | | Upstream response | 32 MiB → 502 `gateway_error`, nothing billed. A 4K Nano Banana 2 image is about 20 MB of base64. | | Generation budget | 450 s, covering the body read as well as the render | | Streaming | not supported | | Admission | one slot held for the whole render (typically 10–60 s) | | Providers | OpenAI direct for the GPT Image models, OpenRouter for the rest — never Azure, Bedrock or an Anthropic deployment | ## Models billed at reported cost The OpenRouter models are billed at **the cost OpenRouter reports** for the render, passed through with no margin ([Pricing & billing](/docs/pricing#models-billed-at-the-providers-reported-cost)). - **Reserve:** `n` × the model's per-image ceiling, in USD — **Per image ≤** on the models page. - **Settle:** the reported cost, exactly. None reported: the ceiling, marked `estimated`. - **The response** keeps OpenAI's shape. `usage.input_tokens` and `usage.output_tokens` are OpenRouter's own counts, for reading only; `usage.cost` is the charge. | Model id | Listed price | Per image ≤ | Measured (2026-09-23) | |---|---|---|---| | `gemini-2.5-flash-image` (Nano Banana) | $30 / M image tokens | $0.05 | $0.0387, 5–7 s | | `gemini-3.1-flash-image` (Nano Banana 2) | $60 / M image tokens | $0.20 | $0.0448 at 512 to $0.1512 at 4K, 7–30 s | | `gemini-3-pro-image` (Nano Banana Pro) | $120 / M image tokens | $0.30 | $0.1344 at 1K and 2K, $0.2409 at 4K, 18–34 s | | `seedream-5.0-pro` | $0.045; $0.09 at 2K | $0.10 | as listed, 40–67 s | | `seedream-5.0-lite` | $0.035 | $0.04 | as listed, 14–30 s | | `grok-imagine-image-2.0` | $0.04–$0.08 by quality and resolution | $0.09 | $0.06 at 1K, $0.08 at 2K (medium), 32–46 s | ## Response headers | Header | When | Meaning | |---|---|---| | `x-request-id` | every call | the id `GET /v1/generation?id=` audits | | `x-tm-provider` | served requests | deployment that rendered the image | | `x-tm-attempts` | after ≥ 1 dispatch | physical dispatches, failovers included | | `x-tm-upstream-status` | when a provider answered | the provider's own HTTP status | | `x-tm-dropped-params` | when non-empty | comma-joined names of parameters the gateway stripped | | `x-tm-upstream-model` | aggregator swaps only | the id the gateway actually sent | | `x-tm-served-by` | aggregator routes that name it | the provider the aggregator used | | `x-tm-route-plan-id` | billed requests | the route plan that chose the deployment | | `x-tm-admission-mode`, `x-tm-remaining-rpm`, `x-tm-remaining-tpm` | billed requests | how the request was admitted and the headroom left — see [Limits and capacity](/docs/admission) | | `x-tm-error-code` | every failure | the canonical error class | | `x-tm-error-origin`, `x-tm-limit-scope`, `x-tm-limit-kind` | limit, provider and infrastructure failures | where it came from, and the limit that refused it | There is no image-specific header. Browsers can read all of these except `x-tm-error-code` and `x-tm-route-plan-id`, which are not CORS-exposed. ## Examples curl — decode the first image straight to a file: ```bash curl -s https://api.routerplus.com/v1/images/generations \ -H "Authorization: Bearer $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"gpt-image-1","prompt":"a red bicycle on a white background","size":"1024x1024","quality":"medium"}' \ | jq -r '.data[0].b64_json' | base64 -d > out.png ``` Python — the official `openai` SDK with the base URL swapped: ```python import base64 import os from openai import OpenAI client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) r = client.images.generate( model="gpt-image-1", prompt="a red bicycle on a white background", size="1024x1024", quality="medium", ) with open("out.png", "wb") as f: f.write(base64.b64decode(r.data[0].b64_json)) print(r.usage.model_dump()["cost"]) # the exact ledger debit, in USD ``` TypeScript — the official `openai` package: ```typescript import { writeFileSync } from "node:fs"; import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY }); const r = await client.images.generate({ model: "gpt-image-1", prompt: "a red bicycle on a white background", size: "1024x1024", quality: "medium", }); const b64 = r.data?.[0]?.b64_json ?? ""; writeFileSync("out.png", Buffer.from(b64, "base64")); console.log(r.usage); // input/output tokens and cost (USD) ``` ## Content-free Prompts and image bytes are content, and the marketplace never stores content. The prompt is never logged and never written to the ledger; the base64 bytes and any `revised_prompt` are relayed to you and dropped. Telemetry keeps the response **byte size** so throughput stays measurable, and nothing else. See [Data policy](/docs/data-policy). ## Model discovery ```bash curl -s "https://api.routerplus.com/v1/models?output_modalities=image" \ -H "Authorization: Bearer $TM_API_KEY" ``` Every row of `GET /v1/models` carries `architecture.output_modalities` — `["text"]`, `["image"]` or `["video"]` — and the `?output_modalities=` filter takes a comma list of those values. The public feed `https://app.routerplus.com/api/models.json` carries `output_modalities`, the full `supported_parameters` descriptor map for image models, and `max_output_tokens` as the per-image ceiling. See [GET /v1/models](/docs/api-models) and [Models & catalog](/docs/models). ## BYOK Image generation works through an `openai` provider connection. List the exact image ids you intend to send on the connection — matching is literal, and an id the connection does not carry is a `404 model_unavailable`. The model's metadata (ceiling, descriptors) still comes from the house catalog. `usage.cost` is `0`, the marketplace charge is zero, and OpenAI bills your own account for the render. See [Bring your own key](/docs/byok). See also: [Models & catalog](/docs/models) · [Pricing & billing](/docs/pricing) · [Errors](/docs/errors) · [Wire compatibility](/docs/compat) --- Source: https://app.routerplus.com/docs/api-videos.md # POST /v1/videos Generate a video from a text prompt. This is the OpenAI Videos API surface: point `client.videos.create` at `https://api.routerplus.com/v1` and the call works unchanged. A video takes from ten seconds to several minutes to make, longer than any edge keeps a request open. So a video is a **job**. `POST /v1/videos` answers at once with the job; you read the job with `GET /v1/videos/{id}` until it is `completed`; then you download the MP4 from `GET /v1/videos/{id}/content`. To try it without writing code, pick a video model in the [playground](/docs/playground). Today every video model is served through OpenRouter, on the marketplace's own account: a key bound to your own provider connection cannot make videos yet. Text-to-video only: image-to-video and reference images are not offered yet, and a request that sends one is refused. ## Authentication | Header | Format | |---|---| | `Authorization` | `Bearer tm_vk_...` (the standard OpenAI-style header) | | `x-api-key` | `tm_vk_...` — also accepted, same key | A missing or invalid key returns 401 with `error_type: "auth"`. ## Create a video `POST /v1/videos` with a JSON body, capped at 1 MB. Every rule below is checked **before** anything is reserved or sent to a provider, so a refused request has no ledger row and no charge. | Parameter | Type | Behavior | |---|---|---| | `model` | string, required | A video model id — see [Model discovery](#model-discovery). A chat model here is a 400 naming the chat routes; an image model is a 400 naming `/v1/images/generations`; an id that is not a video model anywhere is a 404. | | `prompt` | string, required | Non-empty, at most 32 000 characters. | | `seconds` | integer, or a string of digits | The length. OpenAI's SDKs send a string (`"8"`), and both forms are read the same way. Each model accepts a range (below). **Omitted = the model's default length**, which is always sent, so the job is billed for the length that was reserved. | | `duration` | integer | OpenRouter's name for `seconds`. Send one of the two; if you send both and they differ, it is a 400. | | `resolution` | string | Per model, for example `480p`, `720p`, `1080p`. | | `aspect_ratio` | string | Per model, for example `16:9`, `9:16`, `1:1`. | | `size` | string | Exact pixels, for example `1280x720`, on the models that declare sizes (Seedance). Elsewhere it is dropped and recorded. | | `generate_audio` | boolean | On the models that make sound. | | `seed` | integer | On the models that declare one. | | `user` | string | Your end-user id, at most 256 characters. Kept with the job and returned on it; never sent to the provider. | | `provider` | object | Routing controls read by the gateway and never forwarded — `require_parameters` and the rest, see [Routing policies](/docs/routing-policies). | The accepted values are **per model** and published in `/api/models.json` under `supported_parameters`. A value outside a model's declared set is a 400 whose message starts with the field name and lists what that model accepts. An explicit `null` means "not provided", as on the other routes. **Refused, never dropped:** `input_reference`, `frame_images` and `input_references`. To drop an image would make, and bill, a different video from the one asked for. **Dropped and recorded:** any other key, including `callback_url` — the gateway does not send provider webhooks to your URL. The names are in the `x-tm-dropped-params` header and on the ledger row. Send `"provider": {"require_parameters": true}` to turn a drop into a 400 instead. The answer is HTTP 200 with the video object, `status: "queued"`. It carries the same headers as an image response: `x-request-id`, `x-tm-provider`, `x-tm-attempts`, `x-tm-upstream-status`, and `x-tm-dropped-params` when something was dropped. The provider has 60 seconds to accept the job; past that the request is a 504 `upstream_error`. ## The video object ```json { "id": "video_7f3c2a9b1e4d4c6a8b0f2e1d3c5a7b9e", "object": "video", "model": "wan-3.0", "status": "completed", "progress": 100, "created_at": 1790000000, "completed_at": 1790000094, "expires_at": null, "seconds": "5", "size": null, "resolution": "720p", "aspect_ratio": "16:9", "error": null, "usage": { "cost": 0.5 } } ``` | Field | Meaning | |---|---| | `id` | Ours, `video_` and 32 hex digits. It is never the provider's id. | | `status` | `queued` → `in_progress` → `completed` or `failed`. | | `progress` | `0` until the video is done, then `100`. The providers report no progress, and the gateway does not invent one. | | `seconds` | A string, as in OpenAI's object. | | `error` | On a failed job: `{code, message}`. The code is one of `video_generation_failed`, `content_policy_violation`, `video_cancelled`, `video_expired`, `video_timeout`. The provider's own words are never kept. | | `usage.cost` | USD, once the job has ended: the charge. Absent while the job runs, and on static dev keys. | | `user` | Your `user`, when you sent one. | ## Read, list and download | Call | Answer | |---|---| | `GET /v1/videos/{id}` | The video object. Ask every few seconds; a job takes minutes. | | `GET /v1/videos` | Your organization's videos, newest first: `{object: "list", data, first_id, last_id, has_more}`. Query: `limit` (1–100, default 20), `order` (`desc` or `asc`), `after` (a video id). | | `GET /v1/videos/{id}/content` | The MP4, streamed from the provider. `?variant=video` is the only variant (anything else is a 400). A `Range` header is passed on, so a player can seek; a range the provider cannot satisfy is a 416. Before the job is `completed` it is a 409. | The gateway follows a job itself: it asks the provider every 5 seconds for the first two minutes, then every 10 seconds, then every 20. A read between two of those asks shows the last status the gateway knows. A video of another organization is a 404, exactly like a video that does not exist. The file is streamed through as it arrives, with the provider's `content-type` (else `video/mp4`), `content-length`, `content-range` and `accept-ranges`, and is **never stored by the marketplace** — save it when you download it. ## Billing Video models are billed at **the cost the provider reports** for the job (OpenRouter's `usage.cost`), pass-through, with no margin. See [Pricing & billing](/docs/pricing#models-billed-at-the-providers-reported-cost). - **Reserve**, before the job is submitted: `seconds` × the model's per-second ceiling. The models page shows that ceiling as **Per second ≤**. Your balance must cover it. - **Settle**, when the job ends: the reported cost, exactly. The hold is then released. - **A failed, cancelled or expired job costs nothing.** You get no video, so you pay nothing. - **A completed job without a reported cost** is charged its reservation, marked `estimated` in [GET /v1/generation](/docs/api-usage). - **If the gateway restarts while a job runs**, the job is charged its reservation, as for any request a crash leaves open. The job itself continues, and `GET /v1/videos/{id}` still answers for it. A job the gateway follows for two hours without an end is also charged its reservation, and ends `failed` with the code `video_timeout`. Worked numbers on `wan-3.0` (per-second ceiling $0.20): a 5-second request reserves $1.00 (5 × $0.20). At 480p the job reported $0.2125 when measured on 2026-09-23; that is the charge, and the rest of the hold is released. ## Models | Model id | Made by | Length | Resolutions | Sound | Listed price | Per second ≤ | |---|---|---|---|---|---|---| | `seedance-2.5` | ByteDance | 4–30 s (default 5) | 480p, 720p; 12 exact sizes | yes | $10.70 per million video tokens (about $0.10/s at 480p, $0.23/s at 720p) | $0.30 | | `hailuo-3-max` | MiniMax (H3 Max) | 5–15 s (default 5) | 480p, 768p | no | $0.05/s at 480p, $0.08/s at 768p | $0.08 | | `wan-3.0` | Alibaba | 5–30 s (default 5) | 480p, 720p, 1080p | yes | $0.05/s at 480p, $0.10/s at 720p, $0.20/s at 1080p | $0.20 | OpenRouter bills Wan 3.0 for at least 5 seconds, so shorter lengths are not offered: a 2-second video would cost as much as a 5-second one. Measured on 2026-09-23, Wan 3.0 charged 85% of its listed rate (a 5-second 480p video: $0.2125). Times measured on 2026-09-23: H3 Max finished a 5-second video in about 13 seconds; Seedance 2.5 and Wan 3.0 took about two minutes at 480p. A provider's queue can hold a job much longer — one Seedance job at 720p waited more than 15 minutes. ## Errors | Case | Status | `error_type` | Billed | |---|---|---|---| | A field outside the model's accepted values, an image input, a chat or image model on this route | 400 | `invalid_request` | no row | | The model is listed but has no price row | 400 | `model_not_priced` | no row | | Not a video model in the catalog | 404 | `model_unavailable` | no row | | Request body over 1 MB | 413 | `request_too_large` | no row | | Unknown id, or another organization's video | 404 | `not_found` | — | | The file before the job is completed | 409 | `invalid_request` | — | | Key RPM or a shared limit, balance too low for the reservation, or a monthly spend cap | 429 | `rate_limit` / `insufficient_quota` | no row | | The provider refused the submit | the provider's | `upstream_error` | $0 | | The provider did not accept the job within 60 s | 504 | `upstream_error` | $0 | | The provider accepted the job but named no job id, or its answer could not be read | 502 | `upstream_error` | $0 | | Every deployment serving the model is cooling down | 503 | `gateway_error` | nothing dispatched; `retry-after: 5` | | The provider did not send the file on a download | 502 | `upstream_error` | — | | The provider accepted the job but it failed | 200 on the read; `status: "failed"` | — | $0 | Bodies follow the OpenAI error envelope with the canonical class in `error.metadata.error_type`; a provider's own error body is never relayed. The reverse refusal also holds: a video model sent to `POST /v1/chat/completions`, `POST /v1/messages` or `POST /v1/images/generations` is a 400 naming `/v1/videos`. ## Examples curl — create, read, download: ```bash curl -s https://api.routerplus.com/v1/videos \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"hailuo-3-max","prompt":"a paper boat drifting down a rainy street","seconds":"5"}' # → {"id":"video_…","status":"queued",…} curl -s https://api.routerplus.com/v1/videos/video_… -H "Authorization: Bearer $TM_API_KEY" curl -s https://api.routerplus.com/v1/videos/video_…/content -H "Authorization: Bearer $TM_API_KEY" -o video.mp4 ``` Python — the official `openai` SDK with the base URL swapped: ```python import os import time from openai import OpenAI client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) video = client.videos.create(model="wan-3.0", prompt="a lighthouse at dusk", seconds="5") while video.status in ("queued", "in_progress"): time.sleep(5) video = client.videos.retrieve(video.id) if video.status == "completed": client.videos.download_content(video.id).write_to_file("video.mp4") ``` TypeScript — the official `openai` package: ```typescript import { writeFileSync } from "node:fs"; import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY }); let video = await client.videos.create({ model: "seedance-2.5", prompt: "a fox leaping into snow", seconds: "5" }); while (video.status === "queued" || video.status === "in_progress") { await new Promise((r) => setTimeout(r, 5000)); video = await client.videos.retrieve(video.id); } if (video.status === "completed") { const file = await client.videos.downloadContent(video.id); writeFileSync("video.mp4", Buffer.from(await file.arrayBuffer())); } ``` The SDKs' types list OpenAI's own models and lengths; the gateway reads any model id and any whole number of seconds the model accepts. ## Content-free A prompt and a video are content, and the marketplace never stores content. The gateway keeps a job's id, model, length, shape, status and charge; never the prompt, never the provider's error text, never the file. See [Data policy](/docs/data-policy). ## Model discovery ```bash curl -s "https://api.routerplus.com/v1/models?output_modalities=video" \ -H "Authorization: Bearer $TM_API_KEY" ``` Every row of `GET /v1/models` carries `architecture.output_modalities`: `["text"]`, `["image"]` or `["video"]`. The public feed `https://app.routerplus.com/api/models.json` carries `supported_parameters` for each video model, `max_output_tokens` as its per-second ceiling in cost units (one micro-dollar each), and `price_card`, the provider's listed prices. See also: [POST /v1/images/generations](/docs/api-images) · [Pricing & billing](/docs/pricing) · [Errors](/docs/errors) --- Source: https://app.routerplus.com/docs/api-decisions.md # POST /v1/decisions Ask a decision model typed questions about a state, and get one typed answer per question. A decision model does not write text. You send a **state** (a message, a ticket, a record) and your **questions**. Each answer is a choice, a score, a yes/no, an order, a set of tags or extracted fields, with the probability or confidence behind it. Use it for routing, classification, moderation and other decision points in your code. The catalog has ten decision models: | Model | Catalog id | Made by | Providers | Question format | |---|---|---|---|---| | Jev 1.13 | `typesafe/jev-1.13` | TypeSafe | `typesafe`, then `openrouter-decisions` | System One `questions`: `type` is `choice`, `score` or `noul` | | Mercury Decide | `inception/mercury-decide` | Inception | `openrouter-decisions-free` | System One `questions`, exactly as for Jev | | Bespoke Nimble v3 | `bespokelabs/nimble-v3` | Bespoke Labs | `bespokelabs` | System One `questions`, as for Jev, within Nimble's own limits | | Clef | `cloudflare/clef` | Cloudflare | `workers-ai` | System One `questions`, as for Jev, within Clef's own limits | | Clef-flash | `cloudflare/clef-flash` | Cloudflare | `workers-ai` | System One `questions`, as for Jev, within Clef's own limits | | Decider 2B | `routerplus/decider-2b` | RouterPlus | `routerplus` | System One `questions`, exactly as for Jev | | Kev 4B | `routerplus/kev-4b` | RouterPlus | `routerplus` | System One `questions`, exactly as for Jev | | Perplexity Decider v1 27B | `perplexity/pplx-decider-v1-27b` | Perplexity | `perplexity-decisions` | System One `questions`, as for Jev, within its own limits; it bills the state once per question | | Sage 1.2 | `levanto/sage-1.2` | Levanto | `levanto` | System One `questions`, as for Jev, within Sage's own limits; it bills the state once per question | | GLiNER-2.5-Decide | `fastino/gliner-2.5-decide` | Fastino | `fastino` | A GLiNER `schema`: `classifications`, `entities`, `structures`, `relations` | **The question format follows the model.** Every call has the same envelope: `model` and `state` in, `answers` and `usage` out. The questions go in the model's own form: a `questions` map for the System One models, a `schema` object for GLiNER. What goes inside a question, and what comes back inside an answer, is the model's own format. Nine models share one format, **System One** (TypeSafe's): Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage. A body written for one runs on the others with only `model` changed, when it is inside each model's limits. A question in another model's format is a 400. The ten are different models, so the gateway never fails over from one to another. The gateway calls these endpoints: | Model | Order | Provider | Endpoint the gateway calls | Model id sent upstream | |---|---|---|---|---| | Jev | 1 | TypeSafe (direct API) | `POST https://api.typesafe.ai/v1/systemone` | `jev-1.13.0` | | Jev | 2 | OpenRouter | `POST https://openrouter.ai/api/alpha/decisions` | `typesafe/jev-1.13-20260917` | | Mercury Decide | 1 | OpenRouter | `POST https://openrouter.ai/api/alpha/decisions` | `inception/mercury-decide:free` | | Nimble | 1 | Bespoke Labs (direct API) | `POST https://api.bespokelabs.ai/v1/nimble/systemone` | `nimble-v3` | | Clef | 1 | Cloudflare Workers AI (direct API) | `POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef` | `clef` | | Clef-flash | 1 | Cloudflare Workers AI (direct API) | `POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef-flash` | `clef-flash` | | Decider 2B | 1 | RouterPlus (our own service, direct API) | `POST /v1/decisions` on RouterPlus's endpoint | `decider-2b` | | Kev 4B | 1 | RouterPlus (our own service, direct API) | `POST /v1/decisions` on RouterPlus's endpoint | `kev-4b` | | Perplexity Decider | 1 | Perplexity (direct API) | `POST https://api.perplexity.ai/v1/decisions` | `pplx-decider-v1-27b` | | Sage | 1 | Levanto (direct API) | `POST https://sage.levanto.ai/v1/systemone` | `sage-latest` | | GLiNER | 1 | Fastino (direct API) | `POST https://api.fastino.ai/v1/chat/completions` | `fastino/GLiNER-2.5-Decide` | For Jev you send one request shape. The gateway sends each provider the id it expects, and you get one answer shape back, whichever provider served it. If TypeSafe fails before it answers (overloaded, down, rate limited), the gateway tries OpenRouter. Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sage and GLiNER have one provider each, so their requests have no second provider to try. On Workers AI, `{account_id}` is our Cloudflare account, and the path names the model. Decider 2B and Kev 4B share one RouterPlus deployment: one pool and one health circuit serve both. The `x-tm-provider` header names the provider that answered. To compare the models without code, use **Decision** in the [playground](/docs/playground). It takes one set of questions and translates it for each model. ## Authentication | Header | Format | |---|---| | `Authorization` | `Bearer tm_vk_...` | | `x-api-key` | `tm_vk_...` — also accepted, same key | A missing or invalid key returns 401 with `error_type: "auth"`. ## Request JSON body, at most 1 MB (413 `request_too_large` above that). The gateway checks every rule on this page **before** it reserves money or calls a provider. A refused request has no ledger row, no `x-tm-attempts` header and no charge. | Parameter | Type | Behavior | |---|---|---| | `model` | string, required | A decision model id. A dated id of a listed decision model is accepted and routes as the model: `typesafe/jev-1.13-20260917` (OpenRouter's dated id) as `typesafe/jev-1.13`, `inception/mercury-decide-20260930` as `inception/mercury-decide`. OpenRouter's variant suffix (`inception/mercury-decide:free`) is not an id here: send `inception/mercury-decide`. Bespoke's own names (`nimble-v3`, `nimble-latest`) are not ids here either: send `bespokelabs/nimble-v3`. Nor are Cloudflare's (`clef`, `clef-flash`, `@cf/cloudflare/clef`, `@cf/cloudflare/clef-flash`): send `cloudflare/clef` or `cloudflare/clef-flash`. RouterPlus's own short names (`decider-2b`, `kev-4b`) are not ids here: send `routerplus/decider-2b` or `routerplus/kev-4b`. Perplexity's own name (`pplx-decider-v1-27b`) is not an id here either: send `perplexity/pplx-decider-v1-27b`. A chat, image or video model is a 400 that names its route. An unknown id is a 404. | | `state` | required | What the model evaluates. Its forms are the model's: see [System One questions](#system-one-questions) and [GLiNER schema](#gliner-schema) below. | | `questions` | object, the System One models, required | A map of question id to question. You choose the ids; each answer comes back under the same id. At least one question. On GLiNER, `questions` is a 400 that says the model takes a GLiNER `schema`. | | `schema` | object, GLiNER only, required | GLiNER's own schema. See [GLiNER schema](#gliner-schema). On Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider or Sage, `schema` is a 400 that says the model takes `questions`. | | `threshold`, `include_confidence`, `include_spans` | GLiNER only | See [GLiNER options](#gliner-options). | | `provider` | object | Routing controls, read by the gateway and never forwarded: `only`, `order`, `ignore`, `require_parameters` and the rest — see [Routing policies](/docs/routing-policies). The provider ids are `typesafe` and `openrouter-decisions` for Jev, `openrouter-decisions-free` for Mercury Decide, `bespokelabs` for Nimble, `workers-ai` for Clef and Clef-flash (not `cloudflare`, which is an OpenRouter host tag), `routerplus` for Decider 2B and Kev 4B, `perplexity-decisions` for Perplexity Decider (not `perplexity`, which is an OpenRouter host tag), `levanto` for Sage and `fastino` for GLiNER. | ### System One questions Eight models take the same question format, TypeSafe's **System One**: Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B and Perplexity Decider. Everything in this section holds for each of the eight; only `model` differs, and Nimble, Clef, Clef-flash, the two RouterPlus models and Perplexity Decider have some limits of their own (see [Bespoke Nimble v3](#bespoke-nimble-v3), [Clef and Clef-flash](#clef-and-clef-flash), [Decider 2B and Kev 4B](#decider-2b-and-kev-4b) and [Perplexity Decider v1 27B](#perplexity-decider-v1-27b)). The examples show Jev; put another System One model's id in `model`, for example `inception/mercury-decide`, `bespokelabs/nimble-v3`, `cloudflare/clef`, `routerplus/decider-2b` or `perplexity/pplx-decider-v1-27b`, and the same body runs on that model. **State.** A non-empty string, or any JSON object or array (a chat log, a record, application state). Questions can point at a field by name: `` "Is `ticket.body` urgent?" ``. **Questions.** Every question has a `type`, `instructions` and, for most types, `criteria`. `instructions` is a non-empty string, or a JSON object or array: put the question in one field and the data it refers to in others. | `type` | `criteria` | The answer | |---|---|---| | `choice` | Required. An object of option name to description, 1 to 255 options. A description can be a string, JSON, or `null` when the name says enough. | `choice` (the most likely option), `probabilities` (every option), `confidence` | | `score` | Required. An ordered array of level descriptions, low to high, 2 to 10 levels. | `score` (probability-weighted, can fall between levels), `legend`, `probabilities`, `confidence` | | `noul` | Optional. `{"true": "...", "false": "..."}` — what yes and no mean. | `noul`: the probability that the answer is yes, 0 to 1 | A question takes only `type`, `instructions` and `criteria`. Any other field in a question is a 400 that names it. A question with `kind` (the format of Levanto's own `/decide` API) is a 400 that says the model takes System One questions. Ask many questions in one call: the model reads the state once and answers every question against it, so one call with ten questions is cheaper and faster than ten calls. Perplexity Decider and Sage are the exceptions: they read and bill the state once per question, so on them one call saves round trips, not tokens (see [Perplexity Decider v1 27B](#perplexity-decider-v1-27b) and [Sage 1.2](#sage-12)). ```json { "model": "typesafe/jev-1.13", "state": { "ticket": "Help! My payouts have been failing for 3 days." }, "questions": { "team": { "type": "choice", "instructions": "Which team should handle `ticket`?", "criteria": { "billing": "Payments, invoices, refunds", "technical": "Bugs, outages, integrations", "sales": null } }, "frustration": { "type": "score", "instructions": "How frustrated is the customer?", "criteria": ["Calm", "Annoyed", "Angry"] }, "urgent": { "type": "noul", "instructions": "Is this urgent?", "criteria": { "true": "Explicitly time-sensitive", "false": "No urgency expressed" } } } } ``` #### Mercury Decide Inception's Mercury Decide is a System One model: the state, the questions and the answers above are its format too. What is its own: | | Mercury Decide | |---|---| | Catalog id | `inception/mercury-decide`. The dated `inception/mercury-decide-20260930` routes as the same model. | | Route | OpenRouter's decisions endpoint only (`POST https://openrouter.ai/api/alpha/decisions`, provider id `openrouter-decisions-free`), as `inception/mercury-decide:free`. Inception's own API does not serve it yet. | | Price | $0 in and $0 out: OpenRouter serves only the model's free variant today, and the gateway passes that through. `usage.cost` is `0`. See [Billing](#billing). | | Context | 32,768 tokens, the state and every question together. | | Our limit | The deployment's pool takes 20 requests a minute in all, and each organization at most 15 of them; at most 5 requests in flight at once, 3 per organization; and 700,000 tokens a minute in all, 525,000 per organization. Tokens are held at the gateway's estimate of the body while a request runs (a state near the 32k context is held at about 44,000) and corrected to OpenRouter's count when it settles (about 33,000 for that state), so the 15-per-organization share of requests is the limit that binds, not tokens. A burst above any of these is a 429 `rate_limit` from the gateway, with `retry-after` and the `x-tm-limit-*` headers, before anything reaches OpenRouter: `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown` for the pause after a relayed OpenRouter 429, next row) and `x-tm-limit-id` says whether it was the pool (`pool:…`) or your organization's share of it (`pool-share:…`). | | OpenRouter's limit | A free variant shares OpenRouter's daily cap on free requests across our whole account (1,000 a day). When that is used up, OpenRouter answers 429 until its daily reset, and the gateway relays it as 429 `rate_limit` with OpenRouter's `retry-after` and `x-tm-upstream-status: 429`. A relayed 429 pauses the deployment's pool for that `retry-after`, at least 1 s and at most 60 s: a request inside the pause gets the gateway's own 429 `rate_limit`, with `x-tm-limit-kind: cooldown`, `x-tm-limit-id: pool:…` and no `x-tm-upstream-status`. After the pause, requests reach OpenRouter again and meet the cap again, until its daily reset. A relayed 429 never counts against the deployment's health circuit, so the daily cap is 429s until the reset, not the 503 of a cooling-down deployment. The gateway does not count the daily cap itself, so a request can pass our pool and still meet it. | | Retention | OpenRouter publishes no retention terms for this free endpoint, so the deployment declares `prompt_logging: "retained"`. See [Data policy](/docs/data-policy). | A free endpoint is, in OpenRouter's own words, not production-suitable. For a decision that must come back, keep Jev as your System One fallback in your own code: the gateway never fails over between models. #### Bespoke Nimble v3 Bespoke Labs' Nimble is a System One model: the state, the questions and the answers above are its format too. The gateway calls Bespoke's own API. What is its own: | | Bespoke Nimble v3 | |---|---| | Catalog id | `bespokelabs/nimble-v3`. Bespoke's names `nimble-v3` and `nimble-latest` are not ids here. | | Route | Bespoke Labs' API only (`POST https://api.bespokelabs.ai/v1/nimble/systemone`, provider id `bespokelabs`), as `nimble-v3`. | | Price | $0.04 per million input tokens. Output is free. This is Bespoke's own price, with no markup. See [Billing](#billing). | | Questions | 1 to 64 questions per call. A `choice` takes 2 to 255 options: a choice with one option, which Jev takes, is refused. A `score` takes 2 to 10 levels through the gateway, as for every System One model (Nimble itself takes up to 255). The gateway does not check Nimble's own limits: a body above them is Bespoke's 422, relayed as an `upstream_error`, and costs nothing. (The Playground's Decision mode does check them, before a run.) | | Structured content | `instructions` and a `choice` option's description may be a JSON object or array, as on Jev. Bespoke's published schema types them as strings, but Nimble takes them: a probe with both answered 200 (2026-10-02). | | Context | 32,768 tokens for each question's prompt, the state included. Nimble never cuts a prompt: a longer one is refused. Bespoke counts the state and the questions once per call, not once per question. | | Busy | Bespoke answers 503 with `Retry-After` while Nimble starts, 529 when it is busy (Bespoke says: retry after about one second) and 502 when the model fails. Each is a retryable `upstream_error` that counts against the deployment's health circuit, like every 5xx: two in a row open it for 30 s, and in that window every request but one probe gets 503 `gateway_error`, "all deployments cooling down". Retry after a second or two, with backoff. The gateway does not relay Bespoke's `Retry-After` on a 503. See [Errors](#errors). | | Numbers | Full precision: Jev's numbers have 2 decimals, Nimble's are not rounded. A very small probability can come back in exponent form, for example `3.7e-06`. `confidence` is 1 when one option has all the probability and 0 when every option is equally likely. It is not the chance that the answer is right. | | Our limit | The deployment's pool takes 7 requests in flight at once, and each organization at most 5 of them (Bespoke allows our account 8; one is kept for our test environment). It also takes 600 requests a minute (450 per organization) and 15,000,000 tokens a minute (11,250,000 per organization). A burst above any of these is a 429 `rate_limit` from the gateway, with `retry-after` and the `x-tm-limit-*` headers, before anything reaches Bespoke; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown` for the pause after a relayed Bespoke 429, next row). | | Bespoke's limit | Bespoke runs at most 8 requests at once for our whole account. Above that it answers 429 with `Retry-After`. The gateway relays it as 429 `rate_limit` with `x-tm-upstream-status: 429`, and pauses the deployment's pool for that `retry-after`, at least 1 s and at most 60 s. A relayed 429 never counts against the deployment's health circuit. | | Retention | Bespoke keeps request content for up to 30 days, to debug the service and check for abuse, and then deletes it. The deployment declares `prompt_logging: "retained"`. Bespoke does not train on API content unless an organization opts in; ours does not. See [Data policy](/docs/data-policy). | #### Clef and Clef-flash Cloudflare's Clef and Clef-flash are System One models: the state, the questions and the answers above are their format too. Clef-flash is the smaller model (9 billion parameters, against Clef's 27 billion): it is faster and costs less. The gateway calls Cloudflare's own API, Workers AI. What is their own: | | Clef and Clef-flash | |---|---| | Catalog ids | `cloudflare/clef` and `cloudflare/clef-flash`. Cloudflare's names `clef`, `clef-flash`, `@cf/cloudflare/clef` and `@cf/cloudflare/clef-flash` are not ids here. | | Route | Cloudflare Workers AI's REST API only, provider id `workers-ai`: `POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef` as `clef`, and `.../@cf/cloudflare/clef-flash` as `clef-flash`. `{account_id}` is our Cloudflare account. The path picks the model, and the body's `model` must be the same model; the gateway sets both. Cloudflare puts each answer inside its own envelope (`result`, `success`, `errors`, `messages`). The gateway opens it, so you get Jev's answer shape. A 200 whose envelope does not say `success: true` is a 502 `upstream_error` and costs nothing. | | Price | Clef: $0.24 per million input tokens. Clef-flash: $0.09 per million input tokens. Output is free: Cloudflare reports `output_tokens: 0` on every answer. This is Cloudflare's own price, with no markup. See [Billing](#billing). | | Questions | 1 to 64 questions per call. Each question id (a key of `questions`) is 1 to 100 letters, digits, `_`, `.` or `-`: it must match `^[A-Za-z0-9_.-]{1,100}$`. A `choice` takes 2 to 255 options: a choice with one option, which Jev takes, is refused. A `score` takes 2 to 10 levels. The gateway does not check Clef's own limits: a body outside them is Cloudflare's 422 (or 400), relayed as an `upstream_error` with Cloudflare's words, and costs nothing. (The Playground's Decision mode does check them, before a run.) | | Structured content | `instructions`, a `score` level and a `choice` option's description may be a JSON object or array, as on Jev, and an option's description may be `null`. Cloudflare's schema says so. A probe with a JSON object as `instructions` and JSON option descriptions answered 200 (2026-10-02). | | Context | 65,536 tokens, the state and every question together. Before the model runs, Cloudflare estimates the request's tokens at about 4 characters a token. It refuses a request estimated above 65,536 tokens (about 262,000 characters) with a 413, which the gateway returns as 413 `context_overflow`; it costs nothing. Cloudflare's schema says that a long state is cut to fit, but in our test (2026-10-02) a request of 522,286 characters was refused, not cut. | | Images | Not accepted through the gateway. Clef's `images` field, Cloudflare's addition to System One, is dropped and recorded in `x-tm-dropped-params`, never forwarded. With `"provider": {"require_parameters": true}` it is a 400. | | `request_id` | Do not send a top-level `request_id`. The gateway drops and records it, and never forwards it: Workers AI reads that field as a lookup in its queue of async requests, and answers 404. | | Numbers | 4 decimals: Jev's numbers have 2, and Nimble's are not rounded. `confidence` is on Clef's own scale, not on Jev's. Cloudflare says only that it is derived from the probabilities. On one live answer (a choice of three options, the top one at 0.8274), Clef's `confidence` was 0.5654, where Jev's formula gives 0.7411. So a threshold that you tuned on Jev's `confidence` does not transfer to Clef: tune it on Clef, or use the probabilities. `confidence` is not the chance that the answer is right. | | Busy | When Workers AI is busy, Cloudflare answers 429 with its code 3040, "Capacity temporarily exceeded, please try again." The gateway relays it as 429 `rate_limit`, with Cloudflare's words and `x-tm-upstream-status: 429`, and the deployment's pool pauses for Cloudflare's `Retry-After`, at least 1 s and at most 60 s. Without a `Retry-After`, the relayed 429 has no `retry-after` either, and the pool pauses for 1 s. A relayed 429 never counts against the deployment's health circuit. Retry after a second or two, with backoff. A 5xx (for example Cloudflare's 500, "Model execution failed") is a retryable `upstream_error` that counts against the circuit, like every 5xx: two in a row open it for 30 s. A 408, Cloudflare's own timeout, is a retryable 504 `upstream_error` that does not count. See [Errors](#errors). | | Our limit | Clef and Clef-flash share one pool. It takes 200 requests a minute (150 per organization), at most 8 requests at a time (6 per organization) and 2,000,000 tokens a minute (1,500,000 per organization). A burst above any of these is a 429 `rate_limit` from the gateway, with `retry-after` and the `x-tm-limit-*` headers, before anything reaches Cloudflare; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown` for the pause after a relayed Cloudflare 429). | | Cloudflare's limit | Cloudflare publishes 300 requests a minute for Clef's task type, Text Generation. It does not say whether Clef and Clef-flash share that limit, so our pool stays below it. Cloudflare's 429 with code 3036 says that our account's free daily allocation is used up. The gateway answers it with 503 `model_unavailable` and `metadata.retryable: false`: that is our account, never yours. It never counts against the deployment's health circuit. | | Retention | Cloudflare says that it does not read, store or train on requests to Clef, or on their responses. The deployment declares `prompt_logging: "none"`. The gateway calls Workers AI directly, never through Cloudflare's AI Gateway, which logs prompts and responses by default. See [Data policy](/docs/data-policy). | #### Decider 2B and Kev 4B RouterPlus is our own decision-model service. Its two models run on our own GPUs in the United States, and both are System One models: the state, the questions and the answers above are their format too. Each reads English. What is their own: | | Decider 2B | Kev 4B | |---|---|---| | Catalog id | `routerplus/decider-2b` | `routerplus/kev-4b` | | Name at RouterPlus | `decider-2b` | `kev-4b` | | Input, per million tokens | $0 | $0 | | Context | 25,600 tokens; a longer input is cut (see below) | 8,192 tokens; a longer input is refused | What holds for both: | | Decider 2B and Kev 4B | |---|---| | Route | RouterPlus's endpoint only (`POST /v1/decisions` there, provider id `routerplus`). RouterPlus's short names in the table above are not ids here. One deployment serves both models, so one pool, one health circuit and one pause after a 429 cover both. | | Price | Free: input and output are $0, and `usage.cost` is `0` on every call. RouterPlus is our own service, so the price is ours to set, not a provider's price passed through. See [Billing](#billing). | | Questions | The gateway checks the System One rules above. RouterPlus takes every form that Jev takes, on both models: `instructions` as a string or as a JSON object or array, `score` levels as strings or as JSON objects (the answer's `legend` gives the levels as you sent them), and a `choice` option's description as a string, JSON or `null`. In probes on 2026-10-02, each model also answered a `choice` with one option, a `choice` of 255 options, a `score` of 2 and of 10 levels, a yes/no with described `true` and `false`, a JSON object or array as the state, and 200 questions in one request. A question that RouterPlus calls malformed is its 400, relayed with RouterPlus's reason for each question id (see [Errors](#errors)); it costs nothing. | | Long input | Decider 2B reads up to 25,600 tokens and cuts a longer input: it keeps the question and its options first, then the start of the state. It drops the rest and does not return an error. Kev 4B never cuts: above 8,192 tokens it refuses the request with a 400 (see [Errors](#errors)). The gateway bills RouterPlus's `usage.input_tokens`, and RouterPlus counts there only the tokens the model can read: at most the model's context. In a test on 2026-10-02, a Decider 2B request of 30,053 tokens billed 25,600. RouterPlus can refuse a long input in place of cutting it (its `truncate` field), but the gateway drops that field (see [Parameters that do not apply](#parameters-that-do-not-apply)), so through the gateway a long input on Decider 2B is always cut. Put what matters at the start of the state, or use Kev 4B for an input up to 8,192 tokens and Decider 2B up to 25,600. | | Response cache | RouterPlus keeps each answer, with its usage, for 600 s (10 minutes). The cache key is the RouterPlus API key, the model and the exact request body, byte for byte. The gateway sends every request with our one RouterPlus key, so an identical request from any of our buyers within 600 s can get the cached answer: the same `answers`, and the same input count in `usage`. Your response keeps its own `id`. RouterPlus marks it with `usage.cached: true` and the header `X-Cache: HIT` in its own response. A cached answer costs $0: the listing prices cached input at $0, and the gateway bills a cached answer's input tokens at that price. The gateway's response marks it with `usage.cached_input_tokens`, equal to `usage.input_tokens`. RouterPlus skips its cache for a request with `"cache": false` or a `Cache-Control: no-cache` header. The gateway forwards neither, so you cannot skip the cache through the gateway. | | Request id | RouterPlus sends `X-Request-Id` (`req_…`) on every answer. The gateway keeps it with the request's record as the provider's id. Your response's `id` and its `x-request-id` header are the gateway's own. | | Cold start | After a quiet period, the first request starts the model on a GPU. This takes 15 to 20 s. The gateway waits up to 60 s for a decision, so a cold start is a slow 200, not an error. | | Our limit | The deployment's pool takes 11,400 requests a minute, 15,000,000 tokens a minute and 59 requests in flight at once. Each organization may use 75 % of each: 8,550 requests a minute, 11,250,000 tokens a minute, 44 at once. The 59 in flight is the burst limit: a pool has no limit per second, and 59 requests in flight at RouterPlus's typical 0.31 s a call carry about 190 a second, just above the pool's 11,400 a minute; with our test environment's 10 a second, that stays within RouterPlus's own limit (next row). Your organization's own limits bind first: by default 300 requests a minute, 1,000,000 tokens a minute and 8 in flight (see [Rate limits & spend caps](/docs/limits)). A burst above any of these is a 429 `rate_limit` from the gateway, with `retry-after` and the `x-tm-limit-*` headers, before anything reaches RouterPlus; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown` for the pause after a relayed RouterPlus 429, next row). | | RouterPlus's limit | About 200 requests a second for the whole endpoint, shared with our test environment. A cached answer comes back sooner than a computed one, so a burst of repeated requests can still reach it. Above it RouterPlus answers 429. The gateway relays it as 429 `rate_limit` with `x-tm-upstream-status: 429`, and pauses the deployment's pool for that `retry-after`, at least 1 s and at most 60 s. A relayed 429 never counts against the deployment's health circuit. | | Busy or failed | A 5xx from RouterPlus is a retryable `upstream_error` that counts against the deployment's health circuit: two in a row open it for 30 s for both models, and in that window every request but one probe gets 503 `gateway_error`, "all deployments cooling down". Retry after a few seconds, with backoff. A RouterPlus 401, 402 or 403 refuses our key, never yours: you get 503 `model_unavailable`, as from every System One provider. See [Errors](#errors). | | Retention | RouterPlus writes no request logs. Its answer cache keeps each answer and its usage for 600 s, under a key made from our RouterPlus key and the request. The cache is all it keeps, so the deployment declares `prompt_logging: "none"`. Nothing is used for training. See [Data policy](/docs/data-policy). | #### Perplexity Decider v1 27B Perplexity's Decider v1 27B is a System One model: the state, the questions and the answers above are its format too. The gateway calls Perplexity's own API. One thing sets it apart from the rest of the family, and it changes the cost: it reads and bills the state once per question. What is its own: | | Perplexity Decider v1 27B | |---|---| | Catalog id | `perplexity/pplx-decider-v1-27b`. Perplexity's own name `pplx-decider-v1-27b` is not an id here. | | Route | Perplexity's API only (`POST https://api.perplexity.ai/v1/decisions`, provider id `perplexity-decisions`), as `pplx-decider-v1-27b`. The provider id is not `perplexity`, which is an OpenRouter host tag. Perplexity answers in Jev's shape, with no envelope. The gateway keeps Perplexity's `x-request-id` with the request's record as the provider's id. | | Price | $0.04 per million input tokens. Output is free. This is Perplexity's own price, with no markup. See [Billing](#billing). | | The state, once per question | Perplexity runs one prompt for each question, and each prompt holds the whole state. It bills the state in each one. In a test on 2026-10-02, a state of 1,843 tokens billed 1,843 tokens with one question and 9,215 with five. So on this model one call with ten questions saves round trips, not tokens: it costs about what ten calls cost. The gateway holds the state once per question too (see Our limit below). The question ids are not billed. `output_tokens` is one per question, at $0. A `choice` with one option runs no prompt and costs nothing: a request of only such questions comes back with 0 tokens and costs $0. | | Questions | 1 to 128 questions per call. Above 128 is Perplexity's 400, "Each request needs between 1 and 128 questions", relayed as an `upstream_error` with Perplexity's words; it costs nothing. A `choice` takes 1 to 255 options, as on Jev. A `score` takes 2 to 10 levels. A question id is any key the gateway takes. The gateway does not check the 128 (the Playground's Decision mode does, before a run). | | Structured content | `instructions`, a `score` level and a `choice` option's description may be a JSON object or array, as on Jev, and an option's description may be `null`. A probe with a JSON object as `instructions`, JSON option descriptions and JSON score levels answered 200 (2026-10-02). | | Context | 262,144 tokens for each question's prompt: the state and that one question. A request with several questions can bill more than that in all: a state of 52,348 tokens with six questions billed 314,088 tokens (2026-10-02). Above the limit, Perplexity refuses the request with a 400, "Input length (262144) exceeds or equals model's maximum context length (262144)". The gateway returns it as 400 `context_overflow`, with `metadata.retryable: false`; it costs nothing. Perplexity does not cut a long state. | | Image parts | Decision models take text and JSON. An object with `"type": "image_url"`, anywhere in the state or in a question, is a 400 from the gateway before any money is reserved, for example `state[1]: an object with "type": "image_url" is an image part, and model "perplexity/pplx-decider-v1-27b" takes text and JSON only`. Other JSON, a field named `type` with another value included, goes as you sent it. | | Numbers | Full precision, as Nimble's. On a `choice`, `confidence` is Jev's: `(p_max − 1/n) / (1 − 1/n)`, so a threshold that you tuned on Jev's choice confidence transfers. On a `score` it does not. Perplexity's score confidence is `1 − Σ pᵢ·|i − m| / ((1/L)·Σⱼ |j − m|)`, where `m` is the most likely level and `L` the number of levels; Jev measures the second sum about the central level, not about `m`. The two agree only when the most likely level is the central one. On one live answer (three levels at 0.756, 0.193 and 0.051), Perplexity's `confidence` was 0.7051, where Jev's formula gives 0.5577. So tune a score threshold on Perplexity Decider, or use the probabilities. `confidence` is not the chance that the answer is right. | | Busy | Perplexity answers 504 when the model does not answer in about a minute; the body is an HTML page. The gateway returns a retryable 504 `upstream_error` that counts against the deployment's health circuit, like every 5xx: two in a row open it for 30 s, and in that window every request but one probe gets 503 `gateway_error`, "all deployments cooling down". The gateway's own wait for a decision (60 s) also ends in a 504. In Perplexity's tests a few hundred input tokens answered in under 2 s, and a prompt near the context in 23 s. A Perplexity 401, 402 or 403 refuses our key, never yours: you get 503 `model_unavailable`, as from every System One provider. See [Errors](#errors). | | Our limit | The deployment's pool takes 540 requests a minute, at most 8 requests at a time and 15,000,000 tokens a minute. Each organization may use 75 % of each: 405 requests a minute, 6 at a time, 11,250,000 tokens a minute. Your organization's own limits bind first: by default 300 requests a minute, 1,000,000 tokens a minute and 8 in flight (see [Rate limits & spend caps](/docs/limits)). Tokens are held at the gateway's estimate while a request runs, with the state counted once per question, and corrected to Perplexity's count when it settles. So a long state with many questions can be larger than a token limit on its own: such a request is a 429 `rate_limit` with `x-tm-limit-kind: tpm` before anything reaches Perplexity, and it cannot pass as it is. Ask fewer questions per call, or send a shorter state. A burst above any limit is a 429 `rate_limit` from the gateway, with `retry-after` and the `x-tm-limit-*` headers; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown` for the pause after a relayed Perplexity 429, next row). | | Perplexity's limit | Perplexity allows our organization 10 requests a second, over a rolling second, shared with our test environment. Our pools keep to that as a minute budget (600 a minute in all), but they do not count seconds: a burst inside one second can pass 10, and then Perplexity's 429 limits it. It also limits large bursts of tokens, and publishes no number for that. Above a limit Perplexity answers 429, "Request rate limit exceeded, please try again later.", with `Retry-After: 1`. The gateway relays it as 429 `rate_limit` with `x-tm-upstream-status: 429`, and pauses the deployment's pool for that `retry-after`, at least 1 s and at most 60 s. A relayed 429 never counts against the deployment's health circuit. | | Retention | Perplexity says that it does not retain query data sent through its API and does not train on it: its API has zero day retention of prompt data by default. Its compute runs on AWS in North America. The deployment declares `prompt_logging: "none"`. See [Data policy](/docs/data-policy). | #### Sage 1.2 Levanto's Sage is a System One model: the state, the questions and the answers above are its format too. The gateway calls Levanto's System One API. Like Perplexity Decider, it reads and bills the state once per question. What is its own: | | Sage 1.2 | |---|---| | Catalog id | `levanto/sage-1.2`. Levanto's own names (`sage-latest`, `levanto-sage-v1.3`) are not ids here. | | Route | Levanto's API only (`POST https://sage.levanto.ai/v1/systemone`, provider id `levanto`), as `sage-latest`. Levanto answers in Jev's shape. | | Price | $0.05 per million input tokens. Output is free. This is Levanto's own price, with no markup. See [Billing](#billing). | | The state, once per question | Levanto reads the state once for each question and bills it each time. So on this model one call with ten questions saves round trips, not tokens. The gateway holds the state once per question too. | | Questions | A `choice` takes 2 to 120 options and a `score` 2 to 26 levels. Every question gets an answer, never `null`. Levanto refuses a request outside its limits; the gateway relays the refusal with Levanto's words, and it costs nothing. | | Not offered | Levanto's own `/decide` format (`kind`, sort, tags), images, grounding and `reasoning`. A question with `kind` is a 400; `reasoning` and `latency_mode` are dropped and recorded. | | Retention | The deployment declares `prompt_logging: "retained"`. See [Data policy](/docs/data-policy). | ### GLiNER schema GLiNER-2.5-Decide is a small encoder model. It reads one text and fills in a schema: it classifies the text against your labels, and it can extract entities, structured fields and relations from it. **State.** Text only: a non-empty string, or `{"kind": "text", "value": "..."}`. A list, an image or any other state is a 400 that names the field. To judge a record, send it as text, for example the JSON as a string. **Schema.** GLiNER takes no `questions`. It takes `schema`, Fastino's own schema object. The gateway checks it and forwards it as it is. `schema` has only these four keys; send at least one, and each key you send must be non-empty (an empty list or object is a 400 that names the key). A flat list in place of the object is a 400 (Fastino has deprecated it). | Key | Form | Limits | |---|---|---| | `classifications` | A list of `{"task", "labels", "multi_label"?, "top_k"?, "cls_threshold"?}` | 1 to 50 tasks. `task` is 1 to 256 characters, unique in the schema, and not `__proto__`. `labels` is 1 to 100 unique non-empty strings. `multi_label` is a boolean. `top_k` is an integer, 1 or more. `cls_threshold` is a number from 0 to 1. Any other key is a 400 that names it. | | `entities` | A list of entity names, or of `{"name", "description"?}` | 1 to 50 entities. Each name is a non-empty string, unique in the list. | | `structures` | An object of structure name to a list of fields, each `"field::type::description"` | 1 to 50 structures. A name is 1 to 256 characters and not `__proto__`. 1 to 50 non-empty field strings per structure. | | `relations` | A list of relation names, or an object of name to `{"description"?, "threshold"?}` | 1 to 50 unique non-empty names. `threshold` is a number from 0 to 1. A `head` or `tail` key is a 400. | **Labels are plain strings.** A label cannot be an object with a description. To describe a label, put the description in the label itself: `"credit: store credit"`. Descriptions matter: on the same refund ticket, the labels `refund`, `credit`, `deny` chose `refund`, and the same labels with descriptions chose `credit`. The answer carries the label text exactly as you sent it. **The task name is the question.** A classification has no `instructions` field. GLiNER reads the task name, so a task name can be a question: a task named `"Does the customer ask for money back?"` with the labels `["yes", "no"]` is a yes/no question. **Result names share one namespace.** GLiNER puts every result at the top level of its answer: each task under its task name, each structure under its structure name, the entities under `entities` and the relations under `relation_extraction`. Two results with one name would overwrite each other at Fastino without an error. So the gateway refuses a collision with a 400 that names both, for example a task named `entities` in a request that also asks for entities. ```json { "model": "fastino/gliner-2.5-decide", "state": "I was charged twice for my March invoice. Please refund the order from Acme Corp. Signed, Jane Doe.", "schema": { "classifications": [ { "task": "action", "labels": ["refund", "credit", "deny"] }, { "task": "issues", "labels": ["double_charge", "angry_customer", "fraud"], "multi_label": true } ], "entities": ["person", { "name": "organization", "description": "business or institution name" }] } } ``` #### GLiNER options Three optional top-level fields go to Fastino as they are. Leave them out to use Fastino's defaults. | Field | Values | What it does | |---|---|---| | `threshold` | A number from 0 to 1. Default 0.5. | The confidence a result needs to be returned. A task's own `cls_threshold` wins for that task. Lower values return more results; higher values return fewer, surer ones. | | `include_confidence` | Boolean. Default `true`. | With `false`, a single-label answer is the bare label string, and a multi-label answer is a list of label strings. | | `include_spans` | Boolean. Default `true`. | With `true`, each entity carries `start` and `end`: its character offsets in the text. | A single-label task returns its top label only; `top_k` does not add more labels to the answer. When no label reaches the threshold, the task's answer is `null`: set `"cls_threshold": 0` on the task to always get the top label and its confidence. A task with `multi_label: true` returns every label at or above the threshold, so `"cls_threshold": 0` returns every label with its confidence. #### Store Fastino keeps each inference unless the request says `store: false`. **The gateway always sends `store: false`.** A `store` field that you send is dropped and recorded, never forwarded. Our Fastino account also has Zero Data Retention on, so Fastino does not train on your text. See [Data policy](/docs/data-policy). #### Not offered yet These parts of Fastino's API are not reachable through the gateway today: - Fine-tuned GLiNER models (Fastino's training-job ids). - Fastino's own batch and async routes. - Several messages. The gateway sends Fastino one user message: the state. GLiNER does not read a system message. ### Parameters that do not apply Decision models take no sampling or output controls. The gateway forwards no `temperature`, `max_tokens`, `stream`, `seed`, `user` or `metadata` to them, and none offers prompt caching, so `cache_control` has nothing to act on. RouterPlus's answer cache (see [Decider 2B and Kev 4B](#decider-2b-and-kev-4b)) is not prompt caching: it works by itself on a whole identical request, and `cache_control` does not reach it. A top-level field that the model does not take is **dropped and recorded** (D8 §2), never forwarded. RouterPlus's own `truncate` and `cache` fields are dropped in this way: | Model | Top-level fields it takes | Examples of fields dropped | |---|---|---| | Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sage | `model`, `state`, `questions`, `provider` | `temperature`, `reasoning`, `latency_mode`, `images`, `request_id`, `truncate`, `cache` | | GLiNER | `model`, `state`, `schema`, `provider`, `threshold`, `include_confidence`, `include_spans` | `store`, `temperature`, `messages` | The dropped names are in the `x-tm-dropped-params` response header and on the attempt's ledger row. To refuse instead, send `"provider": {"require_parameters": true}`: a would-be drop is then a 400 before any money is reserved. ## Response HTTP 200, `application/json`. Every decision model uses the same envelope: `id`, `object`, `model`, `answers` and `usage`. ### System One answers The shape is the same from both of Jev's providers, from Mercury Decide (with `"model": "inception/mercury-decide"` and `"cost": 0`), from Nimble (with `"model": "bespokelabs/nimble-v3"` and its numbers at full precision), from Clef and Clef-flash (with `"model": "cloudflare/clef"` or `"model": "cloudflare/clef-flash"`, its numbers to 4 decimals and `"output_tokens": 0`), from Decider 2B and Kev 4B (with their catalog ids in `model`) and from Perplexity Decider (with `"model": "perplexity/pplx-decider-v1-27b"`, its numbers at full precision and one output token per question): ```json { "id": "5d0c1c4e-3f7a-4c55-9d1b-2f0e7f6f2a10", "object": "decision", "model": "typesafe/jev-1.13", "answers": { "team": { "type": "choice", "choice": "billing", "probabilities": { "billing": 0.88, "technical": 0.12, "sales": 0 }, "confidence": 0.81 }, "frustration": { "type": "score", "score": 1.05, "legend": { "0": "Calm", "1": "Annoyed", "2": "Angry" }, "probabilities": { "0": 0, "1": 0.95, "2": 0.05 }, "confidence": 0.92 }, "urgent": { "type": "noul", "noul": 0.95 } }, "usage": { "input_tokens": 436, "output_tokens": 71, "cost": 0.000018 } } ``` `answers` holds one typed answer per question, under the ids you sent, exactly as the provider returned them. ### GLiNER answers `answers` is GLiNER's own result, verbatim: the JSON that Fastino returns as a string in `choices[0].message.content`, parsed into an object. The gateway adds nothing and removes nothing. For the request in [GLiNER schema](#gliner-schema): ```json { "id": "0b7e4c2a-6f1d-4a39-9c85-3e2d1f7a8b64", "object": "decision", "model": "fastino/gliner-2.5-decide", "answers": { "action": { "label": "refund", "confidence": 0.9327635765075684 }, "issues": [ { "label": "double_charge", "confidence": 0.872347354888916 } ], "entities": { "person": [ { "text": "Jane Doe", "confidence": 0.99609375, "start": 90, "end": 98 } ], "organization": [ { "text": "Acme Corp", "confidence": 0.98828125, "start": 71, "end": 80 } ] } }, "usage": { "input_tokens": 33, "output_tokens": 91, "cost": 0 } } ``` | You asked | Key in `answers` | Value | |---|---|---| | A single-label task | The task name | `{"label", "confidence"}`: the top label. `null` when no label reaches the threshold (0.5 unless you set one; `"cls_threshold": 0` always returns the top label). With `include_confidence: false`, the bare label string. | | A multi-label task | The task name | `[{"label", "confidence"}, ...]`: every label at or above the threshold, possibly none. With `include_confidence: false`, a list of label strings. | | Entities | `entities` | An object of entity name to `[{"text", "confidence", "start", "end"}, ...]`. `start` and `end` are there when `include_spans` is on. | | A structure | The structure name | A list of records, each field as `{"text", "confidence"}`, for example `{"order": [{"id": {"text": "4411", "confidence": 1.0}}]}`. | | Relations | `relation_extraction` | `{"relations": [...]}`. | - **A confidence is not a calibrated probability.** It is GLiNER's score for that label or span, from 0 to 1. Multi-label confidences are independent and do not sum to 1. Compare them within one model, not with System One's probabilities. - GLiNER has no "not sure" answer. A low confidence is the sign that it is unsure. - The gateway accepts a 200 only when the content parses to a JSON object with every task and structure name you asked, plus `entities` and `relation_extraction` when you asked for them. Anything else is an `upstream_error` and costs nothing. ### Usage and headers | Field | Meaning | |---|---| | `id` | The gateway's request id, also in the `x-request-id` header. Look the request up with `GET /v1/generation?id=`. | | `model` | The catalog id that was billed. The provider's own name for the model is in the `x-tm-upstream-model` header. | | `usage.input_tokens` | Input tokens, billed at the model's input rate. For the System One models and GLiNER, the provider's count (Fastino's `prompt_tokens`); on Decider 2B, at most the model's context; on Perplexity Decider and Sage, the state once per question. | | `usage.output_tokens` | Output tokens, billed at the model's output rate. Every decision model prices output at $0; Sage, Clef and Clef-flash report 0, and Perplexity Decider one per question. For GLiNER, Fastino's `completion_tokens`. | | `usage.cached_input_tokens` | Present only on an answer from RouterPlus's cache (Decider 2B, Kev 4B). The part of `usage.input_tokens` that the cache answered: all of it. These tokens are billed at the cached-input price, $0. | | `usage.cost` | USD, the full debit for the request, after any discount, including any attempt that failed over before this one. `0` on BYOK. Absent on static dev keys. | | `usage.cost_before_discount` | Present only when a discount applied. What the request costs at the list price. | | `usage.discount_percent` | Present only when a discount applied. The percent taken off. | Response headers: `x-request-id`, `x-tm-provider` (`typesafe`, `openrouter-decisions`, `openrouter-decisions-free`, `bespokelabs`, `workers-ai`, `routerplus`, `perplexity-decisions`, `levanto` or `fastino`), `x-tm-served-by` (the provider OpenRouter names, when it names one: `Inception` for Mercury Decide), `x-tm-upstream-model` (the provider's own name for the model, for example `clef` or `clef-flash` from `workers-ai`, `decider-2b` from `routerplus`, `pplx-decider-v1-27b` from `perplexity-decisions`), `x-tm-attempts`, `x-tm-upstream-status`, `x-tm-discount-percent` when a discount applied and, when something was dropped, `x-tm-dropped-params`. ## Billing A decision is metered like any other request: key limits, the reservation, the ledger and `usage.cost` all work as on the chat routes. Every decision model is priced per million tokens. Jev, Nimble, Clef, Clef-flash, Perplexity Decider, Sage and GLiNER charge for input only. Mercury Decide is $0 both ways while OpenRouter serves only its free variant; a paid variant, or Inception serving it directly, would be a new listed price, and this page would say so. Decider 2B and Kev 4B run on RouterPlus, our own service, and they are free: $0 in and out, so `usage.cost` is `0` on every call. Each charge rounds down to the micro-dollar. Perplexity Decider bills the state once per question (see [Perplexity Decider v1 27B](#perplexity-decider-v1-27b)): its cost grows with the number of questions as well as with the state. | Model | Input, per million tokens | Output, per million tokens | What one call costs | |---|---|---|---| | Jev | $0.042 | $0 | A three-question call of about 400 input tokens costs about $0.000017. | | Mercury Decide | $0 | $0 | Nothing. The call is still metered, limited and recorded in `GET /v1/generation`, and it still needs an organization with credit. | | Nimble | $0.04 | $0 | A three-question call of about 340 input tokens costs $0.000013. Bespoke's own price, with no markup. | | Clef | $0.24 | $0 | A call of 400 input tokens costs $0.000096. Cloudflare's own price, with no markup. | | Clef-flash | $0.09 | $0 | A call of 400 input tokens costs $0.000036. Cloudflare's own price, with no markup. | | Decider 2B | $0 | $0 | Nothing, as on Mercury Decide. | | Kev 4B | $0 | $0 | Nothing, as on Mercury Decide. | | Perplexity Decider | $0.04 | $0 | Five questions about a state of 1,843 tokens bill 9,215 input tokens: $0.000368. Perplexity's own price, with no markup. | | Sage | $0.05 | $0 | Three questions about a short ticket read 204 input tokens: $0.00001. Levanto's own price, with no markup. | | GLiNER | $0.03 | $0 | A call of 1,000 input tokens costs $0.00003. A call of 33 input tokens, as in the example in [GLiNER answers](#gliner-answers), rounds down to $0. | If a provider answers 200 without a usage object, the gateway bills its own estimate, the same amount it reserved, and marks the attempt `estimated`. For GLiNER it comes from the size of the text and from the tasks, labels and extraction keys in the schema. For Perplexity Decider and Sage it counts the state once per question. A 200 that does not answer every question, or whose GLiNER content is not the object described in [GLiNER answers](#gliner-answers), is an `upstream_error` and costs nothing. After a provider answers 200, the gateway never sends the same request to a second provider. ### Discount A discount may apply to decision models: to every decision model, or to some of them, each at its own percent. It applies from every key, including the playground's. The list price does not change, and the discount can change or end: read it from each response, not from this page. A discounted model's row in `/api/models.json` carries `discount_percent`, and its page shows the percent beside the struck list price. - `usage.cost` is what you pay, after the discount. - `usage.cost_before_discount` is the list price of the request. - `usage.discount_percent` and the `x-tm-discount-percent` header give the percent. - When no discount applies, these three are absent. A discounted charge is the list charge × (100 − percent) / 100, rounded down to the micro-dollar. At 100 % it is 0, and the request is still metered, limited and recorded in `GET /v1/generation`. The discount is for organizations with credit: a balance of $0 is refused with `429 insufficient_quota` whatever the discount, so a new account must complete verified browser sign-in and buy credits before its first decision. On Mercury Decide, Decider 2B and Kev 4B the list price is already $0, so the discount changes nothing: `usage.cost` and `usage.cost_before_discount` are both `0`. When a discount covers one of them, the percent is still reported, as on every model it covers, so your code can read one shape for all ten. ## Errors Errors use the gateway's normal shape — see [Errors](/docs/errors). | Status | `error.code` | When | |---|---|---| | 400 | `invalid_request` | The body failed a rule above: a question in another model's format, a field a question does not take. The message names the field, for example `questions.team.criteria: at most 255 options`. | | 400 | `invalid_request` | The request does not fit the model's grammar: `schema` on Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider or Sage (the message says the model takes `questions`), or `questions` on GLiNER (the message says the model takes a GLiNER `schema`). | | 400 | `invalid_request` | GLiNER only: a state that is not text, a flat-list schema, a schema key or classification key GLiNER does not take, a limit above, two tasks with one name, or two results with one name (the message names both). | | 400 | `invalid_request` | The model is not a decision model. The message names the route that serves it. | | 400 | `invalid_request` | Perplexity Decider: an object with `"type": "image_url"` anywhere in the state or in a question. Decision models take text and JSON. The message names where it is, for example `state[1]: an object with "type": "image_url" is an image part, and model "perplexity/pplx-decider-v1-27b" takes text and JSON only`. It is refused before any money is reserved. | | 400 | `context_overflow` | Perplexity Decider: one question's prompt (the state and that question) is above 262,144 tokens. Perplexity's 400, relayed with its words, for example `upstream perplexity-decisions returned 400: Input length (262144) exceeds or equals model's maximum context length (262144)`. `metadata.retryable` is `false`: make the state shorter. It costs nothing. | | 404 | `model_unavailable` | No decision model has this id. | | 413 | `request_too_large` | The body is larger than 1 MB. | | 413 | `context_overflow` | Clef and Clef-flash: Cloudflare estimated the request above the model's 65,536 tokens and refused it (its 413, code 5021). The message carries Cloudflare's words, for example `upstream workers-ai returned 413: The estimated number of input and maximum output tokens (130569) exceeded this model context window limit (65536).` `metadata.retryable` is `false`: make the state shorter. It costs nothing. | | 503 | `model_unavailable` | Sage: Levanto refused our account (its 401, 402 or 403, for example when our month's usage is used up); the message says the model is out of capacity at the provider. GLiNER: Fastino refused our account for credit or billing (its 402 or 403). The message says the model is out of capacity at the provider. Nimble: Bespoke refused our account for credit (its 402); the message says the model is out of capacity at the provider, and `metadata.retryable` is `false`. Clef and Clef-flash: Cloudflare refused our account, for our token or plan (its 401 or 403), or because the account's free daily allocation is used up (its 429 with code 3036, until 00:00 UTC). The message says the model is out of capacity at the provider, and `metadata.retryable` is `false`. A 401, 402 or 403 from any System One provider (TypeSafe's or OpenRouter's, on Jev or Mercury Decide, and RouterPlus's, on Decider 2B and Kev 4B, and Perplexity's, on Perplexity Decider, too) reads the same: it is our account, never yours. On a connection with your own provider key, that provider's 401, 402 or 403 is your account, and it keeps its status. Other models are not affected. | | 503 | `gateway_error` | "all deployments cooling down", with `retry-after: 5`: the deployment's health circuit is open. On Sage, GLiNER, Nimble, Clef, Clef-flash, the two RouterPlus models and Perplexity Decider this is also what most requests see while the provider refuses our account: each refusal counts against the provider's circuit, so after two of them only one request in each 30 s reaches the provider and gets the message above, and the rest get this one. Decider 2B and Kev 4B share one deployment and so one circuit: two RouterPlus 5xx in a row open it for both. A provider's 429 never counts, so Mercury Decide's daily cap on OpenRouter, Nimble's limit of 8 requests at once, Cloudflare's busy answer on Clef, RouterPlus's limit of about 200 a second and Perplexity's limit of 10 a second stay 429s (below), and Cloudflare's code 3036 stays the 503 above. | | 422 and other 4xx | `upstream_error` | The provider refused the request as wrong. The message carries the provider's own reason, and `x-tm-upstream-status` its status. The gateway does not send a request the provider called wrong to another provider. A Fastino 404 (an unknown upstream model id: our misconfiguration, never yours) is a 502 `upstream_error`, with `x-tm-upstream-status: 404`. A Nimble request above Nimble's own limits (more than 64 questions, a `choice` with one option, or a question's prompt above 32,768 tokens with the state) is Bespoke's 422, relayed here with Bespoke's words; it costs nothing. A Clef or Clef-flash request outside Clef's own limits (more than 64 questions, a question id that does not match `^[A-Za-z0-9_.-]{1,100}$`, or a `choice` with one option) is Cloudflare's 422 (code 5012) or 400 (code 5006), relayed here with Cloudflare's words, for example `upstream workers-ai returned 422: Request body failed validation: questions: Dictionary should have at most 64 items after validation, not 65`; it costs nothing. A request that RouterPlus refuses (Decider 2B, Kev 4B) is RouterPlus's 400, relayed here with RouterPlus's reason and `x-tm-upstream-status: 400`; it costs nothing. RouterPlus refuses a malformed question (`invalid questions`, then the reason for each question id, for example `score needs criteria: a list of 2 to 10 levels`) and a Kev 4B request above 8,192 tokens (the question's id, then `ContextOverflow: branch too long`). A Perplexity Decider request with more than 128 questions is Perplexity's 400, relayed here with Perplexity's words (`Each request needs between 1 and 128 questions`); it costs nothing. | | 429 | `rate_limit` | Mercury Decide: more than its pool allows (20 requests a minute in all, 15 per organization; 5 in flight at once, 3 per organization; 700,000 tokens a minute, 525,000 per organization), with `retry-after` — the gateway refuses before OpenRouter does. `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown`: the pool's pause after a relayed OpenRouter 429, 1 to 60 s) and `x-tm-limit-id` says whether it was the pool (`pool:…`) or your organization's share (`pool-share:…`). Or OpenRouter's daily cap on free requests is used up (1,000 a day across our account): OpenRouter's 429, relayed with its `retry-after` and `x-tm-upstream-status: 429`, until its daily reset; it never opens the deployment's circuit, and it pauses the pool for at most 60 s. | | 429 | `rate_limit` | Nimble: more than its pool allows (7 in flight at once, 5 per organization; 600 requests a minute, 450 per organization; 15,000,000 tokens a minute, 11,250,000 per organization), with `retry-after`, before anything reaches Bespoke; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown`). Or Bespoke already runs 8 requests for our account: Bespoke's 429, relayed with its `Retry-After` and `x-tm-upstream-status: 429`; it never opens the deployment's circuit, and it pauses the pool for at most 60 s. | | 429 | `rate_limit` | Clef and Clef-flash: more than their shared pool allows (200 requests a minute, 150 per organization; at most 8 requests at a time, 6 per organization; 2,000,000 tokens a minute, 1,500,000 per organization), with `retry-after`, before anything reaches Cloudflare; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown`). Or Workers AI is busy: Cloudflare's 429 with code 3040, relayed with Cloudflare's words (`upstream workers-ai returned 429: Capacity temporarily exceeded, please try again.`), its `Retry-After` when it sends one, and `x-tm-upstream-status: 429`. It never opens the deployment's circuit, and it pauses the pool for at most 60 s (1 s when Cloudflare sends no `Retry-After`). Any other Cloudflare 429 is relayed the same way, except code 3036 (the 503 above). | | 429 | `rate_limit` | Decider 2B and Kev 4B: more than their shared pool allows (11,400 requests a minute, 8,550 per organization; 15,000,000 tokens a minute, 11,250,000 per organization; 59 in flight at once, 44 per organization), or more than your organization's own limits, which bind first (by default 300 requests a minute and 8 in flight), with `retry-after`, before anything reaches RouterPlus; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown`). Or RouterPlus is at its limit of about 200 requests a second for the whole endpoint: RouterPlus's 429, relayed with its `retry-after` and `x-tm-upstream-status: 429`; it never opens the deployment's circuit, and it pauses the pool for at most 60 s, for both models. | | 429 | `rate_limit` | Perplexity Decider: more than its pool allows (540 requests a minute, 405 per organization; at most 8 requests at a time, 6 per organization; 15,000,000 tokens a minute, 11,250,000 per organization), or more than your organization's own limits, which bind first (by default 300 requests a minute, 1,000,000 tokens a minute and 8 in flight), with `retry-after`, before anything reaches Perplexity; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown`). The state is held once per question, so a long state with many questions can be larger than a token limit on its own (`x-tm-limit-kind: tpm`): ask fewer questions per call, or send a shorter state. Or Perplexity is at its limit of 10 requests a second for our organization: Perplexity's 429, relayed with its words (`Request rate limit exceeded, please try again later.`), its `Retry-After` and `x-tm-upstream-status: 429`; it never opens the deployment's circuit, and it pauses the pool for at most 60 s. | | 504 | `upstream_error` | A System One provider answered 408, its own timeout. On Clef and Clef-flash the message is `upstream workers-ai returned 408: Request timeout`, with `x-tm-upstream-status: 408`. Retry after a few seconds. A 408 is not tried on a second provider, and it does not count against the deployment's circuit. The gateway's own wait for a decision (60 s) also ends in a 504 (next row). | | 429, 5xx (and 529) | `rate_limit`, `upstream_error` | Every provider failed. A rate limit, an outage or an overload at one provider is tried on the next one first. Levanto answers 503 while Sage loads: retry after a few seconds. Fastino answers 425 or 503 while GLiNER starts (a cold start), and a cold start can also outlast the gateway's wait for a decision (60 s, a 504): both are a retryable `upstream_error`, so retry after a few seconds. A Fastino 429 is a `rate_limit`. Bespoke answers 503 with `Retry-After` while Nimble starts, 529 when it is busy and 502 when the model fails: each is a retryable `upstream_error` that counts against the deployment's circuit (two in a row open it for 30 s), so retry after a second or two. Cloudflare answers 500 when Clef fails ("Model execution failed"): a retryable `upstream_error` that counts against the circuit in the same way. A RouterPlus 5xx is a retryable `upstream_error` too, and it counts against the one circuit of Decider 2B and Kev 4B: two in a row open it for 30 s for both. A RouterPlus cold start (15 to 20 s) is not an error: it fits inside the gateway's 60 s wait and answers 200. Perplexity answers 504 when its model does not answer in about a minute (an HTML page, so the message has no words of Perplexity's): a retryable `upstream_error` that counts against the circuit in the same way. | | 401, 402, 403 | `auth` | Every provider refused our own account with it. This is never your key's fault; it is tried on the next provider first. For GLiNER a 402 or 403, and for every System One model (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage) a 401, 402 or 403, is the 503 above. | Jev reads at most 64 000 tokens per request through TypeSafe (32 000 for the state plus the longest question), and 32 000 through OpenRouter. Mercury Decide reads at most 32 768 tokens per request. Sage reads at most 32 768 tokens per request, the state and every question together. Nimble reads at most 32,768 tokens for each question's prompt, the state included, and never cuts a prompt. Clef and Clef-flash read at most 65,536 tokens per request, the state and every question together, as Cloudflare estimates them (about 4 characters a token); Cloudflare refuses a longer request with the 413 `context_overflow` above. Decider 2B's context is 25,600 tokens and Kev 4B's 8,192. GLiNER-2.5-Decide's context is 8,192 tokens. A request above a provider's limit is refused by that provider. Decider 2B cuts it instead, and bills only the tokens the model read; Kev 4B refuses it with a 400. See [Decider 2B and Kev 4B](#decider-2b-and-kev-4b). Perplexity Decider reads at most 262,144 tokens for each question's prompt (the state and that question), and never cuts one: Perplexity refuses a longer one with the 400 `context_overflow` above. See [Perplexity Decider v1 27B](#perplexity-decider-v1-27b). ## Examples ### Jev ```bash curl https://api.routerplus.com/v1/decisions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"typesafe/jev-1.13", "state":"Help! My payouts have been failing for 3 days.", "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}' ``` Python — no SDK has a decisions method, so send plain HTTP: ```python import os, httpx r = httpx.post( "https://api.routerplus.com/v1/decisions", headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"}, json={ "model": "typesafe/jev-1.13", "state": "Help! My payouts have been failing for 3 days.", "questions": {"team": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "Payments", "technical": "Bugs"}}}, }, ) answer = r.json()["answers"]["team"] print(answer["choice"], answer["confidence"]) ``` Node: ```js const r = await fetch("https://api.routerplus.com/v1/decisions", { method: "POST", headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" }, body: JSON.stringify({ model: "typesafe/jev-1.13", state: { ticket: "Help! My payouts have been failing for 3 days." }, questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } }, }), }); const { answers, usage } = await r.json(); console.log(answers.urgent.noul, usage.cost); ``` ### Mercury Decide The same System One body as for Jev, under Mercury Decide's id. A 429 `rate_limit` here is either our pool (wait `retry-after`) or OpenRouter's daily cap on free requests (wait for its reset): switch on the class, as [Errors](/docs/errors) says, and read the message. ```bash curl https://api.routerplus.com/v1/decisions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"inception/mercury-decide", "state":"Help! My payouts have been failing for 3 days.", "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}' ``` Python — plain HTTP; the answers read exactly as Jev's: ```python import os, httpx r = httpx.post( "https://api.routerplus.com/v1/decisions", headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"}, json={ "model": "inception/mercury-decide", "state": "Help! My payouts have been failing for 3 days.", "questions": { "team": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "Payments", "technical": "Bugs"}}, "angry": {"type": "score", "instructions": "How angry is the customer?", "criteria": ["Calm", "Mild", "Moderate", "Upset", "Furious"]}, }, }, ) if r.status_code == 429: print("rate limited:", r.json()["error"]["message"], "retry after", r.headers.get("retry-after")) else: answers = r.json()["answers"] print(answers["team"]["choice"], answers["angry"]["score"], r.json()["usage"]["cost"]) # cost is 0 ``` Node: ```js const r = await fetch("https://api.routerplus.com/v1/decisions", { method: "POST", headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" }, body: JSON.stringify({ model: "inception/mercury-decide", state: { ticket: "Help! My payouts have been failing for 3 days." }, questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } }, }), }); const { answers, usage } = await r.json(); console.log(answers.urgent.noul, usage.cost); // 0.95, 0 ``` ### Nimble The same System One body as for Jev, under Nimble's id. Keep to Nimble's limits: 1 to 64 questions, and 2 or more options in a choice. ```bash curl https://api.routerplus.com/v1/decisions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"bespokelabs/nimble-v3", "state":"Help! My payouts have been failing for 3 days.", "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}' ``` Python — plain HTTP; the answers read exactly as Jev's, at full precision: ```python import os, httpx r = httpx.post( "https://api.routerplus.com/v1/decisions", headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"}, json={ "model": "bespokelabs/nimble-v3", "state": "Help! My payouts have been failing for 3 days.", "questions": { "team": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "Payments", "technical": "Bugs"}}, "urgent": {"type": "noul", "instructions": "Is this urgent?"}, }, }, ) answers = r.json()["answers"] print(answers["team"]["choice"], answers["urgent"]["noul"]) # e.g. billing 0.997817283712868 ``` Node: ```js const r = await fetch("https://api.routerplus.com/v1/decisions", { method: "POST", headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" }, body: JSON.stringify({ model: "bespokelabs/nimble-v3", state: { ticket: "Help! My payouts have been failing for 3 days." }, questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } }, }), }); const { answers, usage } = await r.json(); console.log(answers.urgent.noul, usage.input_tokens, usage.cost); ``` ### Clef The same System One body as for Jev, under Clef's or Clef-flash's id. Keep to Clef's limits: 1 to 64 questions, ids of 1 to 100 letters, digits, `_`, `.` or `-`, and 2 or more options in a choice. ```bash curl https://api.routerplus.com/v1/decisions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"cloudflare/clef", "state":"Checkout has been failing for every customer for the last hour.", "questions":{"urgent":{"type":"noul","instructions":"Is this support request urgent?"}}}' ``` Python — plain HTTP; the answers read as Jev's, to 4 decimals: ```python import os, httpx r = httpx.post( "https://api.routerplus.com/v1/decisions", headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"}, json={ "model": "cloudflare/clef", "state": "Checkout has been failing for every customer for the last hour.", "questions": { "urgent": {"type": "noul", "instructions": "Is this support request urgent?"}, "team": {"type": "choice", "instructions": "Which team should handle this request?", "criteria": {"billing": "Payments and invoices", "technical": "Bugs and outages", "sales": "Plans and upgrades"}}, "severity": {"type": "score", "instructions": "How severe is the customer impact?", "criteria": ["No impact", "Minor", "Major", "Critical"]}, }, }, ) answers = r.json()["answers"] team = answers["team"] print(team["choice"], team["probabilities"][team["choice"]], team["confidence"]) # e.g. technical 0.8274 0.5654 ``` Node — Clef-flash, the smaller model: ```js const r = await fetch("https://api.routerplus.com/v1/decisions", { method: "POST", headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" }, body: JSON.stringify({ model: "cloudflare/clef-flash", state: { ticket: "Checkout has been failing for every customer for the last hour." }, questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } }, }), }); const { answers, usage } = await r.json(); console.log(answers.urgent.noul, usage.input_tokens, usage.output_tokens); // output_tokens is always 0 ``` ### RouterPlus models The same System One body as for Jev, under Decider 2B's id. For Kev 4B, change only `model`. A first request after a quiet period can take 15 to 20 s while the model starts, so give your client a timeout above 60 s, the gateway's own wait. ```bash curl https://api.routerplus.com/v1/decisions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"routerplus/decider-2b", "state":"Help! My payouts have been failing for 3 days.", "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}' ``` Python — plain HTTP, with a timeout that covers a cold start: ```python import os, httpx r = httpx.post( "https://api.routerplus.com/v1/decisions", headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"}, json={ "model": "routerplus/decider-2b", "state": "Help! My payouts have been failing for 3 days.", "questions": { "team": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "Payments", "technical": "Bugs"}}, "urgent": {"type": "noul", "instructions": "Is this urgent?"}, }, }, timeout=70, ) answers = r.json()["answers"] print(answers["team"]["choice"], answers["urgent"]["noul"], r.headers["x-tm-upstream-model"]) # e.g. billing 0.97 decider-2b ``` Node: ```js const r = await fetch("https://api.routerplus.com/v1/decisions", { method: "POST", headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" }, body: JSON.stringify({ model: "routerplus/decider-2b", state: { ticket: "Help! My payouts have been failing for 3 days." }, questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } }, }), signal: AbortSignal.timeout(70_000), }); const { answers, usage } = await r.json(); console.log(answers.urgent.noul, usage.input_tokens, usage.cost); ``` ### Perplexity The same System One body as for Jev, under Perplexity Decider's id. Keep to its limits: 1 to 128 questions, and text or JSON only. Each question is billed with the whole state, so ask only the questions you need. ```bash curl https://api.routerplus.com/v1/decisions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"perplexity/pplx-decider-v1-27b", "state":"Help! My payouts have been failing for 3 days.", "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}' ``` Python — plain HTTP; the answers read as Jev's, at full precision: ```python import os, httpx r = httpx.post( "https://api.routerplus.com/v1/decisions", headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"}, json={ "model": "perplexity/pplx-decider-v1-27b", "state": "Help! My payouts have been failing for 3 days.", "questions": {"urgent": {"type": "noul", "instructions": "Is this urgent?"}}, }, ) body = r.json() print(body["answers"]["urgent"]["noul"], body["usage"]["input_tokens"]) # e.g. 0.9297849172276498 95 ``` Node: ```js const r = await fetch("https://api.routerplus.com/v1/decisions", { method: "POST", headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" }, body: JSON.stringify({ model: "perplexity/pplx-decider-v1-27b", state: { ticket: "Help! My payouts have been failing for 3 days." }, questions: { team: { type: "choice", instructions: "Which team should handle `ticket`?", criteria: { billing: "Payments", technical: "Bugs" } }, urgent: { type: "noul", instructions: "Is `ticket` urgent?" }, }, }), }); const { answers, usage } = await r.json(); // Two questions: Perplexity bills the state twice. console.log(answers.team.choice, answers.urgent.noul, usage.input_tokens, usage.cost); ``` ### Sage The same System One body as for Jev, under Sage's id. Each question is billed with the whole state, so ask only the questions you need. ```bash curl https://api.routerplus.com/v1/decisions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"levanto/sage-1.2", "state":"Help! My payouts have been failing for 3 days.", "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}' ``` ### GLiNER ```bash curl https://api.routerplus.com/v1/decisions \ -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \ -d '{"model":"fastino/gliner-2.5-decide", "state":"Help! My payouts have been failing for 3 days.", "schema":{"classifications":[{"task":"Is this urgent?","labels":["yes","no"]}]}}' ``` Python — plain HTTP; each answer sits under its task name: ```python import os, httpx r = httpx.post( "https://api.routerplus.com/v1/decisions", headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"}, json={ "model": "fastino/gliner-2.5-decide", "state": "Help! My payouts have been failing for 3 days.", "schema": {"classifications": [{"task": "team", "labels": ["billing: payments", "technical: bugs"]}]}, }, ) answer = r.json()["answers"]["team"] print(answer["label"], answer["confidence"]) # a confidence, not a calibrated probability ``` Node: ```js const r = await fetch("https://api.routerplus.com/v1/decisions", { method: "POST", headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" }, body: JSON.stringify({ model: "fastino/gliner-2.5-decide", state: "Help! My payouts have been failing for 3 days.", schema: { classifications: [ { task: "topics", labels: ["billing", "outage", "fraud"], multi_label: true, cls_threshold: 0 }, ], }, }), }); const { answers, usage } = await r.json(); console.log(answers.topics, usage.input_tokens, usage.output_tokens, usage.cost); ``` ## Model discovery ```bash curl -s "https://api.routerplus.com/v1/models?output_modalities=decisions" \ -H "Authorization: Bearer $TM_API_KEY" ``` Decision models carry `architecture.output_modalities: ["decisions"]` in `GET /v1/models`, and `output_modalities: ["decisions"]` in the public feed `https://app.routerplus.com/api/models.json`. The Anthropic shape of `GET /v1/models` never lists them. --- Source: https://app.routerplus.com/docs/api-models.md # GET /v1/models Lists every model id this gateway can serve — one row per distinct id. This is the endpoint SDK `models.list()` calls resolve to, so it answers in **two shapes**: | Request | Response shape | |---|---| | No `anthropic-version` header | OpenAI list object | | `anthropic-version` header present (any value) | Anthropic models list | The Anthropic SDK sends `anthropic-version` on every request, so each SDK gets its native shape with zero configuration. Authentication is required, in either surface's convention — both work on every gateway route: ``` Authorization: Bearer $TM_API_KEY # or x-api-key: $TM_API_KEY ``` A missing or invalid key is a 401 with `error_type: "auth"` (see [Errors](/docs/errors)). A key bound to your own provider connection lists that connection's models; a key under a routing policy lists the policy's models. Deployed endpoints from [Optimize](/docs/model-search) (`tm/-v`) are not listed here. Your organization's active [dedicated endpoints](/docs/dedicated-endpoints) (`/`) are, with `owned_by` `routerplus`, on the base URL they are served from. ## OpenAI shape (default) ```bash curl -s https://api.routerplus.com/v1/models \ -H "Authorization: Bearer $TM_API_KEY" ``` ```json { "object": "list", "data": [ { "id": "claude-haiku-4-5", "object": "model", "created": 1756900000, "owned_by": "anthropic", "architecture": { "input_modalities": ["text", "image", "file"], "output_modalities": ["text"] } }, { "id": "gpt-image-1", "object": "model", "created": 1756900000, "owned_by": "openai", "architecture": { "input_modalities": ["text"], "output_modalities": ["image"] } }, { "id": "wan-3.0", "object": "model", "created": 1756900000, "owned_by": "openrouter-media", "architecture": { "input_modalities": ["text"], "output_modalities": ["video"] } } ] } ``` | Field | Type | Meaning | |---|---|---| | `id` | string | The exact string for your request's `model` field | | `object` | `"model"` | Constant | | `created` | integer | Unix seconds — see the note below | | `owned_by` | string | Provider id of the first deployment routing would try for this model; `routerplus` for a dedicated endpoint | | `architecture.input_modalities` | array | What a message to the model may carry: `"text"`, plus each of `"image"`, `"file"`, `"audio"` and `"video"` that a route of this key takes. A chat request with a part that no route takes is a 400 (see [Wire compatibility](/docs/compat#content-is-never-silently-dropped)). Image, video and decision models take `["text"]`. A route on your own provider key declares no inputs, so it adds none here; the gateway sends it every part as you sent it. A documented extension OpenAI SDKs ignore | | `architecture.output_modalities` | array | `["text"]`, `["image"]`, `["video"]` or `["decisions"]` — a documented extension OpenAI SDKs ignore | ## Filtering by modality ```bash curl -s "https://api.routerplus.com/v1/models?output_modalities=image" \ -H "Authorization: Bearer $TM_API_KEY" ``` `output_modalities` takes a comma-separated list of `text`, `image`, `video` and `decisions`; the list is a filter, so `?output_modalities=text,image,video,decisions` is the same as sending nothing. Any other value is a 400 `invalid_request` ("invalid output_modalities filter"). The Anthropic shape never lists image, video or decision models — the Anthropic Messages API has no route to call them on. A `provider` query parameter takes the same JSON object as the request-level `provider` routing controls (`{"only":["openai"]}`, for example) and lists what those controls would let through — see [Routing policies](/docs/routing-policies). ## Anthropic shape ```bash curl -s https://api.routerplus.com/v1/models \ -H "x-api-key: $TM_API_KEY" \ -H "anthropic-version: 2023-06-01" ``` ```json { "data": [ { "type": "model", "id": "claude-haiku-4-5", "display_name": "claude-haiku-4-5", "created_at": "2026-09-03T12:00:00.000Z" } ], "has_more": false, "first_id": "claude-haiku-4-5", "last_id": "gpt-4o-mini" } ``` `has_more` is always `false`: the catalog is small and the full list arrives in one response. There is no pagination to implement. > [!NOTE] > `created` / `created_at` is the time the gateway process loaded its catalog, > not the model's release date. Don't build logic on it. For real metadata — > context length, prices, retention policy — use the public > [`/api/models.json`](https://app.routerplus.com/api/models.json), documented in > [Models & catalog](/docs/models). This endpoint stays SDK-shaped plus one > documented extension: `architecture`, which SDKs ignore and modality-aware > clients read. ## SDK usage ```python import os from openai import OpenAI client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"]) for m in client.models.list(): print(m.id, m.owned_by) ``` ```python import os import anthropic # The SDK sends anthropic-version itself, so this gets the Anthropic shape. client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"]) for m in client.models.list(): print(m.id) ``` ```typescript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY }); for await (const m of client.models.list()) { console.log(m.id, m.owned_by); } ``` ## Unknown models on completion calls This list is the contract: a completion request (`/v1/chat/completions` or `/v1/messages`) for any id **not** in it fails immediately with a 404 that echoes exactly the string you sent — before any provider is contacted and before anything is billed. The one exception is a dated pin of a listed model (`claude-haiku-4-5-20251001`, `gpt-4o-2024-08-06`, `-latest`), which routes as its family — see [Models & catalog](/docs/models#date-pinned-ids). OpenAI surface: ```json { "error": { "code": "model_unavailable", "message": "model \"gpt-5o\" is not in the catalog; GET /v1/models lists what this key can serve", "type": "invalid_request_error", "metadata": { "error_type": "model_unavailable" } }, "request_id": "6f3c…" } ``` Anthropic surface: ```json { "type": "error", "error": { "type": "not_found_error", "message": "model \"gpt-5o\" is not in the catalog; GET /v1/models lists what this key can serve", "error_type": "model_unavailable" }, "request_id": "6f3c…" } ``` Both carry the `x-tm-error-code: model_unavailable` header alongside `x-request-id`. > [!WARNING] > There is no fuzzy matching and no silent fallback to a "close" model. A typo > in the model id is a 404, not a quietly different bill. Echoing your own > requested string back names the problem without turning the error into an > existence oracle — unknown ids all 404 identically. Three neighbors of this error are worth telling apart: | Status | `error_type` | Meaning | |---|---|---| | 404 | `model_unavailable` | The id is not in the catalog. Fix the id. | | 503 | `gateway_error` | The id **is** in the catalog, but every deployment serving it is cooling down after failures. Comes with `retry-after: 5` and `x-tm-limit-kind: health` — retry, don't fix. | | 400 | `model_not_priced` | Billed key, and the ledger has no price row for this model — the gateway refuses to serve what it cannot meter. | The full taxonomy is in [Errors](/docs/errors). One more neighbor: a listed **image** model sent to a chat route is a 400 `invalid_request` naming `POST /v1/images/generations`, not a 404 — the id exists, you called the wrong route. A video model on a chat route names `POST /v1/videos` the same way. The mirror case (a chat model on the images or videos route) is the same 400 pointing back at the chat routes. See [POST /v1/images/generations](/docs/api-images) and [POST /v1/videos](/docs/api-videos). --- Source: https://app.routerplus.com/docs/api-usage.md # GET /v1/usage & /v1/generation Two read-only endpoints on the gateway answer "what have I spent?" and "what exactly did this request cost, attempt by attempt?". They read the same ledger your balance is derived from — there is no separate reporting pipeline that could disagree with your bill. Both are **metadata only**: token counts, outcomes, and costs. Prompts and completions are never stored, so they cannot be returned. ## Authentication Both endpoints require a billed key (the first key shown after verified sign-in, or one created in the console) and are scoped to that key's org, workspace and principal. Send it either way: ```bash curl -s https://api.routerplus.com/v1/usage -H "Authorization: Bearer $TM_API_KEY" curl -s https://api.routerplus.com/v1/usage -H "x-api-key: $TM_API_KEY" ``` > [!NOTE] > A static development key (a gateway running without the money plane) has no > ledger to read, so on these routes it gets the same undifferentiated 404 as > any unknown path. ## GET /v1/usage The 20 most recent attempts made with the keys of your workspace and principal, newest first. The org balance fields are filled for a key in the organization's default workspace and default principal — every first key shown after verified sign-in or made in the console — and `null` for a key bound to an explicit workspace or principal (see [Workspaces and principals](/docs/identity)). ```bash curl -s https://api.routerplus.com/v1/usage -H "Authorization: Bearer $TM_API_KEY" | jq ``` ```json { "balance_usd": 0.98971, "credited_usd": 1.0, "spent_usd": 0.01029, "recent_attempts": [ { "request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33", "deployment": "provider-b", "model": "claude-sonnet-4-5", "outcome": "completed", "usage_provenance": "observed", "input_tokens": 1200, "output_tokens": 350, "cost_usd": 0.00669, "billing_source": "house", "inference_cost_usd": 0.00669, "inference_price_source": "platform_list", "price_snapshot_id": "5d2c…", "fee_policy_version": "house-pass-through-v1", "at": "2026-09-04T09:12:33.104Z" }, { "request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33", "deployment": "provider-a", "model": "claude-sonnet-4-5", "outcome": "failed", "usage_provenance": "observed", "input_tokens": 1200, "output_tokens": 0, "cost_usd": 0.0036, "billing_source": "house", "inference_cost_usd": 0.0036, "inference_price_source": "platform_list", "price_snapshot_id": "5d2c…", "fee_policy_version": "house-pass-through-v1", "at": "2026-09-04T09:12:31.512Z" } ] } ``` | Field | Type | Meaning | |---|---|---| | `balance_usd` | number \| null | `credited_usd − spent_usd`, derived from ledger rows on read — never a cached counter | | `credited_usd` | number \| null | sum of all credit grants to the org | | `spent_usd` | number \| null | sum of every settled attempt | | `recent_attempts[].request_id` | string | the request's `x-request-id`; feed it to `/v1/generation` | | `recent_attempts[].deployment` | string | which deployment served the attempt — matches the `x-tm-provider` response header; `routerplus` on a closed [dedicated endpoint](/docs/dedicated-endpoints#closed-endpoints) and on any endpoint's dedicated capacity | | `recent_attempts[].model` | string | the model id you requested; on a closed dedicated endpoint, the endpoint id | | `recent_attempts[].outcome` | string | see the [outcome table](#attempt-outcomes) | | `recent_attempts[].usage_provenance` | string \| null | `observed`, `estimated` or `unknown`; `null` before settlement | | `recent_attempts[].input_tokens` / `output_tokens` | number | provider-reported counts (input is cache-inclusive). On a model billed at the provider's reported cost, `output_tokens` is the charge in micro-dollars | | `recent_attempts[].cost_usd` | number \| null | settled cost; `null` for the brief window before settlement flushes (batched on a ~100 ms cadence) | | `recent_attempts[].at` | string | dispatch timestamp | Both audit endpoints also return the following per-attempt accounting fields: | Field | Meaning | |---|---| | `billing_source` | `house` for marketplace-funded inference; `byok` for inference on your own provider connection | | `inference_cost_usd` | Cost attributed from pinned rates and usage; `null` when either is unknown | | `inference_price_source` | `platform_list`, `customer_rate`, `unknown`, or `legacy_platform_list` for historical records | | `price_snapshot_id` | Identifier of the pinned rate snapshot; `null` when unknown or unavailable historically | | `fee_policy_version` | Economic policy applied to this attempt: `house-pass-through-v1` or `byok-free-v1` | `cost_usd` is the marketplace charge. `inference_cost_usd` is separate cost attribution: it is not a provider invoice, and a list-price or estimated-usage calculation can differ from the customer's actual provider bill. BYOK attempts have a zero marketplace charge and an unknown (`null`) inference cost. Spend caps apply to marketplace charges. See [Bring your own key](/docs/byok) and [routing policies](/docs/routing-policies). Rows are **attempts**, not requests: a request that failed over appears once per physical dispatch, sharing one `request_id`. ## GET /v1/generation?id= The full per-attempt audit for one request. The `id` is the `x-request-id` header every response carries (also `request_id` in `/v1/usage` rows). | Parameter | In | Required | Meaning | |---|---|---|---| | `id` | query | yes | the request id to audit | An `id` that is not a request id at all (not a UUID) is a `404 not_found`. An unknown id — or one belonging to another org, workspace or principal — returns `404` with `"attempts": []` and `"admission_events": []`. Here is a request that failed over once (illustrative deployment ids; yours match your `x-tm-provider` headers): ```bash curl -s "https://api.routerplus.com/v1/generation?id=0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33" \ -H "Authorization: Bearer $TM_API_KEY" | jq ``` ```json { "request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33", "admission_events": [], "attempts": [ { "attempt": 1, "deployment": "provider-a", "model": "claude-sonnet-4-5", "outcome": "failed", "usage_provenance": "observed", "tokens": { "input": 1200, "cached": 0, "cache_write": 0, "output": 0, "reasoning": 0 }, "reserved_max_usd": 0.009, "cost_usd": 0.0036, "billing_source": "house", "inference_cost_usd": 0.0036, "inference_price_source": "platform_list", "price_snapshot_id": "5d2c…", "fee_policy_version": "house-pass-through-v1", "error_code": "upstream_error", "error_origin": "upstream", "admission_context": { "…": "…" }, "route_context": { "…": "…" }, "dropped_params": [], "dispatched_at": "2026-09-04T09:12:31.512Z" }, { "attempt": 2, "deployment": "provider-b", "model": "claude-sonnet-4-5", "outcome": "completed", "usage_provenance": "observed", "tokens": { "input": 1200, "cached": 800, "cache_write": 0, "output": 350, "reasoning": 0 }, "reserved_max_usd": 0.009, "cost_usd": 0.00669, "billing_source": "house", "inference_cost_usd": 0.00669, "inference_price_source": "platform_list", "price_snapshot_id": "5d2c…", "fee_policy_version": "house-pass-through-v1", "error_code": "", "error_origin": null, "admission_context": { "…": "…" }, "route_context": { "…": "…" }, "dropped_params": [], "dispatched_at": "2026-09-04T09:12:33.104Z" } ] } ``` | Field | Type | Meaning | |---|---|---| | `attempt` | number | physical dispatch ordinal, starting at 1 | | `deployment` | string | deployment that served this attempt; `routerplus` on a closed dedicated endpoint | | `outcome` | string | see below | | `usage_provenance` | string | `observed` (provider-reported tokens, or a completed attempt), `unknown` (no usage ever seen), or `estimated` — the provider reported nothing and the gateway priced the attempt itself: on [`/v1/images/generations`](/docs/api-images) a 200 without usage bills `n` × the model's per-image ceiling, the amount the reservation held; on a chat attempt that streamed output, that output at ~4 characters per token | | `tokens` | object | `input` (cache-inclusive), `cached` (reads), `cache_write`, `output`, `reasoning` (subset of output). On a model billed at the provider's reported cost, `output` is the charge in micro-dollars | | `reserved_max_usd` | number | the worst-case hold taken **before** dispatch (ceiling math) | | `cost_usd` | number \| null | what actually settled (floor math); `null` only pre-settlement | | `error_code` | string | canonical error class on failure, empty on success | | `error_origin` | string \| null | where a failure came from — `upstream`, `gateway_admission`, `upstream_quota`, `gateway_infrastructure` or `authorization`; `null` on success | | `admission_context` | object \| null | the limits this attempt claimed, as recorded at dispatch — see [Limits and capacity](/docs/admission) | | `route_context` | object \| null | the route plan behind this attempt: plan id, the deployments planned and excluded, the policy revisions in force; on a dedicated endpoint, `dedicated` (the endpoint, its revision and the role) | | `dropped_params` | array | the parameter paths the gateway dropped for this deployment — the same names as the `x-tm-dropped-params` header | On a closed [dedicated endpoint](/docs/dedicated-endpoints#closed-endpoints), `deployment` and `provider` are `routerplus`, `model` is the endpoint id, `route_context` keeps only `dedicated`, `admission_context` is null, `dropped_params` is empty, and `admission_events` carry no pool `scope_id`. On any dedicated endpoint, its dedicated capacity reads as `routerplus`. `admission_events` lists the limit or balance refusals recorded for this request before any attempt (`origin`, `scope`, `scope_id`, `limit_kind`, `reason`, `status`, `occurred_at`); it is empty when nothing was refused. `reserved_max_usd` versus `cost_usd` is the reserve-then-settle contract made visible: the hold is deliberately pessimistic and is released in full; the settle is computed from provider-reported usage and rounds down. Details in [Pricing & billing](/docs/pricing). ### Attempt outcomes | `outcome` | Meaning | Billing | |---|---|---| | `completed` | full response delivered | settled on reported usage | | `failed` | attempt failed; a later attempt may have served you | settled on any usage the provider reported before dying (often 0) | | `cancelled` | you disconnected mid-request | settled on usage streamed up to the disconnect; if you disconnected before the attempt left the gateway, it settles at zero with provenance `observed`. On `/v1/images/generations` the render is never aborted: it finishes upstream and the attempt settles the provider's reported usage | | `incomplete` | provider died after your stream committed; you received a terminal error event. On `/v1/videos`, a job the gateway gave up following after two hours | settled on usage observed so far; a video job that timed out settles at its reservation | | `unknown_after_crash` | the gateway could not observe the outcome (it stopped mid-attempt; the attempt is recovered from the gateway's durable ledger journal at restart, or by the operator reconciler) | settles **at the reservation** — the call may have been billed upstream, so the ledger keeps the conservative number | > [!WARNING] > You are debited for **every settled attempt** of a request, including ones > discarded by pre-commit failover — an upstream that dies after reading your > prompt has still billed prompt tokens. That is why the wire's `usage.cost` > is the request total, and why this endpoint exists: the slices are never > hidden. In the example above the wire reported `"cost": 0.01029` — > `0.0036 + 0.00669`, both attempts, no averaging. ## Recomputing your bill from the wire Sum of attempt costs equals the response's `usage.cost`: ```bash curl -s "https://api.routerplus.com/v1/generation?id=$REQ" -H "Authorization: Bearer $TM_API_KEY" \ | jq '[.attempts[].cost_usd] | add' ``` And each attempt's cost recomputes from its token counts and the published prices — the exact settle math, floor division and all: ```python import os import requests from decimal import Decimal GATEWAY = "https://api.routerplus.com" H = {"Authorization": f"Bearer {os.environ['TM_API_KEY']}"} prices = {m["id"]: m["pricing_usd_per_million"] for m in requests.get("https://app.routerplus.com/api/models.json").json()["models"]} def per_m(p, sku, fallback): # absent SKUs bill at the base rate — never $0 return int(Decimal(p.get(sku, p[fallback])) * 1_000_000) # micro-USD per M def settle_micro(t, p): non_cached = max(0, t["input"] - t["cached"] - t["cache_write"]) reasoning = min(max(0, t["reasoning"]), t["output"]) plain_out = t["output"] - reasoning total = (non_cached * per_m(p, "prompt", "prompt") + t["cached"] * per_m(p, "cached_prompt", "prompt") + t["cache_write"] * per_m(p, "cache_write", "prompt") + plain_out * per_m(p, "completion", "completion") + reasoning * per_m(p, "internal_reasoning", "completion")) return total // 1_000_000 # floor: settlement rounds down gen = requests.get(f"{GATEWAY}/v1/generation", params={"id": os.environ["REQ"]}, headers=H).json() micro = sum(settle_micro(a["tokens"], prices[a["model"]]) for a in gen["attempts"] if a["outcome"] != "unknown_after_crash") print(micro / 1e6) # equals the wire's usage.cost ``` (`unknown_after_crash` attempts are excluded because they settle at the reservation, not from token counts — compare their `reserved_max_usd`. The same math holds for a model billed at the provider's reported cost: its prices are `"0"` in and `"1"` out per million, and `output` is the charge in micro-dollars. An attempt with `inference_price_source` `customer_rate` was billed at a [dedicated endpoint's](/docs/dedicated-endpoints#prices) contract rate card: recompute it with those prices instead.) One request re-priced at today's catalog can differ from what you were charged only if the price changed since dispatch: attempts settle against **frozen price snapshots** pinned when the request was reserved, which is the recomputation guarantee's fine print — see [Pricing & billing](/docs/pricing). ## Related - [Pricing & billing](/docs/pricing) — the reserve/settle math these numbers come from - [Errors & remediation](/docs/errors) — what each `error_code` means and what to do - [Quickstart](/docs/quickstart) — signup to first billed call ## Who can read what Usage and generation reads are restricted to the authenticated workspace and principal, including that principal's other keys. A request made under another workspace or principal is a 404, like one that never existed. Keys bound to an explicit workspace or principal receive `null` for the org balance, credit and spend totals; a key in the default workspace and principal gets them. Members of the organization see the org-wide accounting in the console at [https://app.routerplus.com/console/usage](https://app.routerplus.com/console/usage) and [https://app.routerplus.com/console/logs](https://app.routerplus.com/console/logs). --- Source: https://app.routerplus.com/docs/api-signup.md # Retired signup API `POST https://app.routerplus.com/v1/signup` has been removed and returns HTTP `404`. Public account creation is available through Clerk in the browser. The former local magic-link submission route, `POST https://app.routerplus.com/login`, also returns HTTP `404`. Open [Sign up](https://app.routerplus.com/signup), complete Clerk authentication and email verification, then save the first API key shown after sign-in. You can create additional keys at [API keys](https://app.routerplus.com/console/keys). Each raw key is shown once. Existing customers should [sign in](https://app.routerplus.com/login) with the same verified email to retain their account, balance and keys. Signup does not grant free credit. Standard paid credit purchases start at **$10**. We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match. See [Pricing](/docs/pricing) and [Billing](https://app.routerplus.com/console/billing) for the offer available to your account. Existing `tm_vk_` API keys and the gateway's authentication headers are unchanged. Continue with [Quickstart](/docs/quickstart) or [Authentication](/docs/authentication). --- Source: https://app.routerplus.com/docs/migration.md # Migrate Most migrations are three changes: the base URL, the API key, and — if you are coming from OpenRouter — the model id. The gateway speaks both major wire dialects, so your SDK stays: - OpenAI-compatible: `POST https://api.routerplus.com/v1/chat/completions` - Anthropic-compatible: `POST https://api.routerplus.com/v1/messages` Every chat model is callable from both surfaces (image models use [POST /v1/images/generations](/docs/api-images)) — the gateway translates requests, streams, and errors in both directions. The exact rules are on [Wire compatibility](/docs/compat). > [!NOTE] > No key yet? Open [Sign up](https://app.routerplus.com/signup), complete Clerk authentication > and email verification, and save the first key shown after sign-in. Buy credits > in [Billing](https://app.routerplus.com/console/billing); signup does not add free credit. Existing customers can > [sign in](https://app.routerplus.com/login) with the same verified email and create a key at > [API keys](https://app.routerplus.com/console/keys). Browser signup replaces the retired > programmatic signup route. Auth is forgiving on purpose: every route accepts the key as either `Authorization: Bearer $TM_API_KEY` or `x-api-key: $TM_API_KEY`, so whichever header your SDK sends, it works. ## From the OpenAI SDK Two lines. `gpt-*` model ids are the same bare ids the provider uses (`gpt-4o-mini`, `gpt-4.1-mini`, `gpt-4o`), so model strings usually survive untouched. ```diff from openai import OpenAI client = OpenAI( - api_key=os.environ["OPENAI_API_KEY"], + base_url="https://api.routerplus.com/v1", + api_key=os.environ["TM_API_KEY"], ) ``` ```diff import OpenAI from "openai"; const client = new OpenAI({ - apiKey: process.env.OPENAI_API_KEY, + baseURL: "https://api.routerplus.com/v1", + apiKey: process.env.TM_API_KEY, }); ``` The same client now reaches Claude models too — no second SDK: ```python client.chat.completions.create( model="claude-sonnet-4-5", # Anthropic model, OpenAI wire — the gateway translates stream=True, messages=[{"role": "user", "content": "hello"}], ) ``` > [!NOTE] > When a request omits `max_tokens` (or `max_completion_tokens`), the gateway > writes the default of 4096 into it before dispatch, on every route. Send > `max_tokens` explicitly if you want a different ceiling; 32,768 is the most a > request may ask for. **Codex CLI** — current releases do not work: since February 2026 Codex calls only the OpenAI Responses API, which the gateway does not implement (`POST /v1/responses` is a 404). A Codex release from before February 2026 takes a provider block in `~/.codex/config.toml`: ```toml [model_providers.tm] name = "RouterPlus" base_url = "https://api.routerplus.com/v1" env_key = "TM_API_KEY" wire_api = "chat" [profiles.tm] model_provider = "tm" model = "gpt-4o-mini" ``` Then run `TM_API_KEY=$TM_API_KEY codex --profile tm`. ## From the Anthropic SDK Same shape: base URL plus key. The Anthropic-compatible surface lives at the gateway root (no `/v1` suffix in `base_url` — the SDK appends `/v1/messages` itself). ```diff from anthropic import Anthropic client = Anthropic( - api_key=os.environ["ANTHROPIC_API_KEY"], + base_url="https://api.routerplus.com", + api_key=os.environ["TM_API_KEY"], ) ``` Claude model ids match Anthropic's own (`claude-sonnet-4-5`, `claude-haiku-4-5`, dated snapshots included) — and GPT models are callable from this SDK too, over the same `/v1/messages` wire. **Claude Code** — environment variables only, no config file: ```bash ANTHROPIC_BASE_URL=https://api.routerplus.com ANTHROPIC_AUTH_TOKEN=$TM_API_KEY claude ``` Your `anthropic-beta` header is forwarded when an Anthropic-dialect provider serves the request, filtered to an allowlist of betas that do not change what a token costs (`prompt-caching`, `claude-code`, `interleaved-thinking` and the like — the full list is on [Authentication](/docs/authentication)); a dropped value is named in `x-tm-dropped-params`. Claude Code's `?beta=true` query string is tolerated and routes normally (query strings are not forwarded upstream). ## From OpenRouter The wire is deliberately close: same OpenAI-compatible endpoint shape, usage in the final stream chunk, a `cost` field in `usage`, and the same optional `HTTP-Referer` / `X-Title` attribution headers. ```diff client = OpenAI( - base_url="https://openrouter.ai/api/v1", - api_key=os.environ["OPENROUTER_API_KEY"], + base_url="https://api.routerplus.com/v1", + api_key=os.environ["TM_API_KEY"], ) completion = client.chat.completions.create( - model="anthropic/claude-sonnet-4.5", + model="claude-sonnet-4-5", stream=True, messages=[{"role": "user", "content": "hello"}], ) ``` What changes, honestly: | OpenRouter | Here | |---|---| | `author/model` slugs (`anthropic/claude-sonnet-4.5`) | Bare provider ids (`claude-sonnet-4-5`) — the full list is [https://app.routerplus.com/api/models.json](https://app.routerplus.com/api/models.json) | | Variant suffixes (`:free`, `:floor`, `:nitro`) | None. A suffixed id is just an unlisted id and returns 404 | | `usage.cost` in credits | `usage.cost` in USD, computed with the exact integer math the ledger settles with, covering the full request debit — failed-over attempts included | | `GET /api/v1/generation?id=` for post-hoc usage | `GET https://api.routerplus.com/v1/generation?id=REQUEST_ID` — returns every physical attempt with per-attempt tokens, outcome, and settled cost | | Out of credits → HTTP 402 | HTTP 429 `insufficient_quota` — byte-compatible with what OpenAI's own SDKs already handle | | `route`, the `models` fallback array, `transforms` | Dropped on every route, and the drop is named in the `x-tm-dropped-params` header. Remove them | | The `provider` object (`sort`, `max_price`, `allow_fallbacks`, …) | A different shape here: `only`, `ignore`, `order`, `upstream`, `allow_fallbacks`, `require_parameters` and the connection controls on [Routing policies](/docs/routing-policies). Any other key inside `provider` — `sort`, `max_price` — is a 400. Remove or translate it | > [!WARNING] > There is no silent model substitution here, ever. An id not in the catalog is > an honest 404 naming what you asked for — that is a design guarantee, not a > limitation. Check every migrated model string against > `https://app.routerplus.com/api/models.json` before you ship. Failover exists, but it never switches providers mid-answer: if a provider dies after output started, you get the tokens streamed so far plus one clean terminal error event, and retry semantics stay yours. Details on [Wire compatibility](/docs/compat). ## Migrate with one prompt Paste this into your coding agent and let it do the whole thing: ```text Migrate this project to RouterPlus. 1. If TM_API_KEY is not set, have me open https://app.routerplus.com/signup and complete Clerk authentication and email verification in the browser. Existing customers sign in at https://app.routerplus.com/login and create a key at https://app.routerplus.com/console/keys. Save the key as TM_API_KEY (env var and .env; never commit it). Continue once the key is available; there is no programmatic signup endpoint. 2. Find every LLM client in this codebase and repoint it: - OpenAI SDK / any OpenAI-compatible client: base_url -> "https://api.routerplus.com/v1", api_key -> TM_API_KEY - Anthropic SDK: base_url -> "https://api.routerplus.com", api key -> TM_API_KEY - Claude Code: run with ANTHROPIC_BASE_URL=https://api.routerplus.com ANTHROPIC_AUTH_TOKEN=$TM_API_KEY - Codex CLI (only a release from before February 2026; current Codex calls only the Responses API, which the gateway does not implement): add a [model_providers.tm] block with base_url "https://api.routerplus.com/v1", env_key "TM_API_KEY", wire_api "chat", and a [profiles.tm] using it. - OpenRouter clients: also strip "author/" prefixes and ":variant" suffixes from model ids, and remove route/models/transforms request fields. Inside a "provider" object keep only: only, ignore, order, upstream, allow_fallbacks, require_parameters; remove the rest (sort, max_price, ...). - Raw HTTP: change the host to https://api.routerplus.com and the bearer/x-api-key to TM_API_KEY. 3. Check every model id against https://app.routerplus.com/api/models.json and substitute the closest listed id where needed. Never invent an id: unlisted models return 404. 4. Run one real streamed test call per repointed client and show me each response's final usage object (it must contain "cost"). 5. Keep the diff minimal. Do not refactor unrelated code. ``` ## Verify the cutover One streamed call proves the whole path — auth, routing, billing: ```bash curl -N https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"claude-haiku-4-5","stream":true,"max_tokens":40,"messages":[{"role":"user","content":"Reply with exactly: MIGRATED"}]}' ``` The final usage chunk must carry `cost` (USD). Then confirm the ledger saw it: ```bash curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usage ``` Every served response also carries an `x-tm-provider` header naming the deployment that served it and `x-tm-attempts` counting physical dispatches — useful when verifying that traffic really moved. If anything fails, every error body has a stable `error_type`; the remediation table is on [Errors](/docs/errors). --- Source: https://app.routerplus.com/docs/install.md # Coding agents These docs are markdown-first on purpose: the primary reader is often a coding agent. Every page is dual-served — `/docs/` rendered for humans, `/docs/.md` as raw markdown for agents — and the machine index lives at [https://app.routerplus.com/llms.txt](https://app.routerplus.com/llms.txt). This page covers wiring **Claude Code** and **Codex CLI** through the gateway and an **installation runbook** for a coding agent. The user completes signup and email verification through Clerk in a browser before the agent configures an API client. To build the gateway into a product instead, give your agent the [Agent integration guide](agent-integration.md). ## Claude Code Claude Code speaks the Anthropic wire format, which the gateway's `/v1/messages` surface accepts directly. Two environment variables, no other change: ```bash export ANTHROPIC_BASE_URL=https://api.routerplus.com export ANTHROPIC_AUTH_TOKEN=$TM_API_KEY claude ``` Or as a one-shot launch: ```bash ANTHROPIC_BASE_URL=https://api.routerplus.com ANTHROPIC_AUTH_TOKEN=$TM_API_KEY claude ``` > [!WARNING] > If this machine's Claude Code is **signed in** (claude.ai subscription), the > login session outranks env keys: interactively you'll get a one-time "use > this API key?" prompt — approve it — but in headless `-p` mode there is no > prompt and requests fail with a 401. For headless/CI use, point Claude Code > at a fresh config dir so the env key is the only credential: > `CLAUDE_CONFIG_DIR=$(mktemp -d) ANTHROPIC_BASE_URL=https://api.routerplus.com ANTHROPIC_API_KEY=$TM_API_KEY claude -p "..."` What makes this work: - `ANTHROPIC_AUTH_TOKEN` is sent as `Authorization: Bearer …`, which the gateway accepts everywhere (as it does `x-api-key`). - Claude Code appends query strings (`/v1/messages?beta=true`); the gateway routes on the pathname, so these pass through cleanly. - Model ids resolve against the catalog at [https://app.routerplus.com/api/models.json](https://app.routerplus.com/api/models.json), which lists one clean id per model (`claude-haiku-4-5`, `claude-sonnet-4-5`, …). Claude Code often sends **dated** ids (`claude-haiku-4-5-20251001`); those resolve to their family automatically and your original id goes to the provider verbatim — see [Models & catalog](/docs/models). A genuinely unknown id gets an honest `404` naming it. - Every request Claude Code makes lands in `https://api.routerplus.com/v1/usage` with its token counts and `cost_usd`, so you can watch what a session costs. > [!NOTE] > The gateway serves these buyer endpoints — `/v1/chat/completions`, > `/v1/messages`, `/v1/messages/count_tokens` (unbilled, models with an > Anthropic-dialect deployment), `/v1/models`, `/v1/images/generations`, > `/v1/videos`, `/v1/usage`, `/v1/generation`, `/v1/route` and `/v1/limits` — > plus an unauthenticated `/healthz`. Anything else an authenticated tool probes > receives a well-formed `404` with `error_type: not_found` in the caller's own > dialect, not a hang or an HTML page. ### Persistent setup Exports vanish with the shell. To make the gateway Claude Code's default, put the env block in `~/.claude/settings.json` (or a project's `.claude/settings.json` to scope it to one repo): ```json { "env": { "ANTHROPIC_BASE_URL": "https://api.routerplus.com", "ANTHROPIC_AUTH_TOKEN": "tm_vk_..." } } ``` ### Pinning which models Claude Code uses Claude Code picks its own model ids; steer them with the standard variables, using any Claude id from the catalog: ```bash export ANTHROPIC_MODEL=claude-sonnet-5 # main model export ANTHROPIC_SMALL_FAST_MODEL=claude-haiku-4-5 # background/fast tasks ``` Inside a session, `/model` accepts any catalog id the same way. Spend from every session lands in [the console](https://app.routerplus.com/console/usage) and `https://api.routerplus.com/v1/usage`, per request. ## Codex CLI > [!WARNING] > Current Codex CLI releases do not work with the gateway. Since February 2026, Codex > calls only the OpenAI Responses API > ([openai/codex discussion #7782](https://github.com/openai/codex/discussions/7782)), > and a config with `wire_api = "chat"` stops with an error. The gateway does not > implement the Responses API: `POST /v1/responses` returns a `404` with > `error_type: not_found`. Codex releases from before February 2026 still speak Chat Completions. For those, pin the wire in the provider block: ```toml # ~/.codex/config.toml [model_providers.tm] name = "RouterPlus" base_url = "https://api.routerplus.com/v1" env_key = "TM_API_KEY" wire_api = "chat" [profiles.tm] model_provider = "tm" model = "gpt-4o-mini" ``` ```bash TM_API_KEY=tm_vk_... codex --profile tm ``` On such a release, `wire_api = "chat"` is a requirement, not a preference: a profile left on the Responses wire fails on every call. ## Other OpenAI-compatible harnesses Most coding tools that speak "OpenAI-compatible" take the same two values: **base URL** `https://api.routerplus.com/v1` and **API key** `tm_vk_...`. The pattern, in the tools' own vocabulary: | Tool | Where | |---|---| | Cursor | Settings → Models → API Keys: set the OpenAI key to your `tm_vk_...` key and enable "Override OpenAI Base URL" with `https://api.routerplus.com/v1` | | Cline / Roo | Provider: "OpenAI Compatible" → Base URL `https://api.routerplus.com/v1`, key `tm_vk_...`, model id from the catalog | | Continue | `models` entry with `provider: "openai"`, `apiBase: "https://api.routerplus.com/v1"`, `apiKey: "tm_vk_..."` | | aider | `OPENAI_API_BASE=https://api.routerplus.com/v1 OPENAI_API_KEY=tm_vk_... aider --model openai/claude-sonnet-5` | | Anything else | If it has a "base URL" box next to its OpenAI key box, those two values are the whole integration | Every chat model in the catalog works through this surface — Claude models included (image models use [POST /v1/images/generations](/docs/api-images), video models [POST /v1/videos](/docs/api-videos)); the gateway translates. Two honest caveats: - Tools built on the OpenAI **Responses** API (the OpenAI Agents SDK's default, and every current Codex CLI release) do not work yet: `POST /v1/responses` returns a 404. Pick the tool's Chat Completions mode where one exists. - These are configuration patterns, not per-version walkthroughs — tool UIs move. The two values above are the invariant. ## SDKs inside the project your agent is editing | Client | Change | |---|---| | OpenAI SDK (any language) | `base_url = "https://api.routerplus.com/v1"`, `api_key = TM_API_KEY` | | Anthropic SDK (any language) | `base_url = "https://api.routerplus.com"`, `api_key = TM_API_KEY` | | Raw HTTP | host → `https://api.routerplus.com`, key in `Authorization: Bearer` or `x-api-key` | Full streaming examples for both surfaces are in the [Quickstart](/docs/quickstart); for migrating a whole codebase in one prompt, see [Migrate](/docs/migration). ## The installation runbook This section is written to be executed. Give it to your agent — `read https://app.routerplus.com/docs/install.md and get me set up` — or walk it yourself. **Objective:** this environment makes a successful, streamed, billed model call through the gateway. **Done when:** a streamed completion ends with a usage object containing a `cost` field, and `https://api.routerplus.com/v1/usage` shows the request. 1. **Get a key** (skip if `TM_API_KEY` is already set). The user opens [Sign up](https://app.routerplus.com/signup), completes Clerk authentication and email verification, and saves the first key shown after sign-in. Signup does not grant free credit: add paid credits in [Billing](https://app.routerplus.com/console/billing). Existing customers sign in at [https://app.routerplus.com/login](https://app.routerplus.com/login) with the same verified email and create a key at [API keys](https://app.routerplus.com/console/keys). Raw keys are shown once. Account creation has no terminal API; do not call the retired `POST /v1/signup` or local `POST /login` routes. Continue client setup when the key is available. 2. **Export the key**, and write it to the project's `.env` if one exists (never commit it): ```bash export TM_API_KEY=tm_vk_... ``` 3. **Discover the models.** Any id in this file works in step 4: ```bash curl -s https://app.routerplus.com/api/models.json ``` 4. **Make the test call.** OpenAI-compatible surface; works for every chat model, including Claude models — the gateway translates: ```bash curl -N https://api.routerplus.com/v1/chat/completions \ -H "Authorization: Bearer $TM_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"claude-haiku-4-5","stream":true,"max_tokens":40,"messages":[{"role":"user","content":"Reply with exactly: MARKETPLACE OK"}]}' ``` Expected: an SSE stream ending with a usage chunk that has a `cost` field, then `data: [DONE]`. On HTTP `429` with `error_type: insufficient_quota`, stop and have the user check their balance and spend caps in [Billing](https://app.routerplus.com/console/billing). Retry **once** after they resolve the cause — never in a loop. The full retry discipline per error class is in [Errors](/docs/errors). 5. **Wire the user's stack.** Apply every branch that matches this environment: the OpenAI SDK, Anthropic SDK, and raw-HTTP changes from the table above; Claude Code exactly as configured earlier on this page, and Codex CLI only when its release is from before February 2026 (see [Codex CLI](#codex-cli)). 6. **Verify billing** — the done-when condition: ```bash curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usage ``` The test call must appear in `recent_attempts` with its token counts and `cost_usd`. 7. **Optional, recommended:** add attribution headers to the user's app — `HTTP-Referer: https://the-users-app.example` and `X-Title: The App Name`. Content-free; they identify the app for future per-app analytics and expose no request data. Finally, report to the user: the org id, where the key is stored, which stacks were wired, and the cost of the test call. ## /llms.txt — start here if you are an agent ```bash curl -s https://app.routerplus.com/llms.txt ``` It carries the gateway base URL, browser signup instructions, links to every doc page in raw-markdown form, and the machine-readable model catalog. Fetch it first; everything else on this page is reachable from it. Its first link is the [Agent integration guide](https://app.routerplus.com/docs/agent-integration.md), the one page to read before you build the gateway into a product. To load every page with one fetch, use [https://app.routerplus.com/llms-full.txt](https://app.routerplus.com/llms-full.txt). ## Signup requires the browser Public signup and sign-in use Clerk at [Sign up](https://app.routerplus.com/signup) and [Sign in](https://app.routerplus.com/login). Complete the authentication and email-verification steps there. Local magic-link submission and programmatic signup are retired; there is no `dev_verify_link` in a public signup response. After signup, agents can use the existing `tm_vk_` key with the gateway normally. ## Next steps - [Quickstart](/docs/quickstart) — the same flow by hand, with streaming examples on both surfaces - [Migrate](/docs/migration) — the one-prompt migration for an existing codebase - [Errors](/docs/errors) — every `error_type`, and the retry discipline agents should follow - [Rate limits & spend caps](/docs/limits) — trial RPM, monthly caps, and the 429 with a reset time --- Source: https://app.routerplus.com/docs/compat.md # Wire compatibility This page is the published form of decision D8 — the conventions your bill is computed from. Changing anything here is treated internally as a billing-contract change and requires a recorded decision. ## Two surfaces, one catalog | Surface | Endpoint | Auth | |---|---|---| | OpenAI-compatible | `POST https://api.routerplus.com/v1/chat/completions` | `Authorization: Bearer` or `x-api-key` | | Anthropic-compatible | `POST https://api.routerplus.com/v1/messages` | `Authorization: Bearer` or `x-api-key` | | OpenAI Images | `POST https://api.routerplus.com/v1/images/generations` | same | | OpenAI Videos | `POST https://api.routerplus.com/v1/videos` | same | Every cataloged chat model is callable from both chat surfaces; image models answer on the Images surface only, video models on the Videos surface only. `GET https://api.routerplus.com/v1/models` lists what your key can serve — OpenAI list shape by default, Anthropic shape when you send an `anthropic-version` header. One distinction drives everything below: - **Passthrough route** — your dialect matches the serving provider's (OpenAI format → OpenAI-dialect provider, Anthropic format → Anthropic-dialect provider). The parameters your dialect defines are forwarded as you sent them, and the gateway adds `stream_options.include_usage` on streamed OpenAI-dialect dispatches. Whatever the provider accepts for those parameters, you get — including native features like Anthropic prompt caching (`cache_control`), `thinking`, `top_k`, and the `anthropic-beta` values listed below. A top-level key outside your dialect's parameter set is dropped and recorded, never forwarded. Session ids and cache marks follow [Sessions and prompt caching](#sessions-and-prompt-caching). - **Translated route** — dialects differ, and the gateway translates the request, the stream, and errors. Translation covers a pinned parameter set; everything else follows the rules below. A "typed 400" in the translation tables is decided per deployment. When a translation refuses your request but another deployment of the model speaks your own dialect, the request goes there instead; you get the 400 only when no deployment can serve it. Content a model does not take follows the same rule: a deployment that does not take a part is skipped, and the 400 comes only when no deployment takes it (see [Content is never silently dropped](#content-is-never-silently-dropped)). ## Parameter handling on passthrough routes The forwarded set per dialect. Anything else at the top level is a recorded drop (see the note under the translation tables). | Dialect | Forwarded as sent | |---|---| | OpenAI → OpenAI-dialect provider | `model`, `messages`, `stream`, `stream_options`, `temperature`, `top_p`, `max_tokens`, `max_completion_tokens`, `stop`, `n`, `tools`, `tool_choice`, `parallel_tool_calls`, `response_format`, `reasoning_effort`, `seed`, `user`, `logprobs`, `top_logprobs`, `frequency_penalty`, `presence_penalty`, `logit_bias`, `metadata`, `store`, `prediction` | | Anthropic → Anthropic-dialect provider | `model`, `messages`, `system`, `max_tokens`, `stream`, `temperature`, `top_p`, `top_k`, `stop_sequences`, `tools`, `tool_choice`, `metadata`, `thinking`, `output_config`, `cache_control` | Some OpenAI-format keys go only to the providers that define them: - `prompt_cache_retention` and `safety_identifier` go to OpenAI and Azure. - The top-level `cache_control` (automatic caching) goes to OpenRouter. - A `cache_control` mark on a content part passes unchanged to every OpenAI-dialect provider. OpenRouter uses it; OpenAI and Azure accept it and ignore it. On Bedrock, the top-level `cache_control` becomes a mark on the last block that can carry one, because Bedrock's API does not take the top-level field. On every route, a top-level `cache_control` that Anthropic would refuse is dropped and recorded as `cache_control`: a field other than `{"type": "ephemeral"}` with an optional `ttl` of `"5m"` or `"1h"`, a fifth mark, a 1-hour field after a 5-minute mark, or a field whose TTL differs from the mark on the last block. Two Anthropic parameters are refused on every route, because stripping them would change who runs what: `mcp_servers` and `container` are a typed 400. The `anthropic-beta` header is forwarded to Anthropic-dialect providers only for values that do not change how a token is billed: `prompt-caching`, `token-efficient-tools`, `fine-grained-tool-streaming`, `interleaved-thinking`, `claude-code`, `oauth` and `computer-use` prefixes. Any other value is dropped and recorded as `header:anthropic-beta:`. ## Parameter handling on translated routes ### OpenAI surface → Anthropic-dialect provider | Parameter | Handling | |---|---| | `model`, `messages`, `stream`, `temperature`, `top_p` | Translated / copied verbatim | | `max_tokens`, `max_completion_tokens` | Translated. The gateway sets `max_tokens: 4096` on any request that carries neither, before dispatch; see [POST /v1/chat/completions](/docs/api-chat-completions) for the 1–32,768 bound | | `stop` (string or array) | → `stop_sequences` | | `system` / `developer` messages | → the Anthropic `system` string, joined in order with blank lines. If a part has a `cache_control` mark, `system` becomes text blocks with the same text, and each mark ends a block | | `cache_control` on a text part | Kept on the Anthropic block. A `role:"tool"` message's mark goes on its `tool_result` | | Other keys of a text part | Not forwarded, and not recorded | | `cache_control` (top-level) | → the Anthropic top-level field (automatic caching). On Bedrock it becomes a mark on the last block | | `tools`, `tool_choice` | Translated: `auto`→`{type:"auto"}`, `required`→`{type:"any"}`, `{function:{name}}`→`{type:"tool",name}`, `none`→`{type:"none"}`. With `none` on a tool-free history, tools are omitted entirely so you don't pay for definitions you forbade using. A tool without `parameters` gets an empty object schema | | `role:"tool"` messages | → `tool_result` blocks; consecutive results merge into one Anthropic user message | | An assistant message with `content: null` and no `tool_calls` | Dropped and recorded (OpenAI's own refusal shape); it has nothing Anthropic can carry | | `stream_options` | Usage is on by default; an explicit `stream_options.include_usage: false` is honored — you are still billed, but no usage chunk is emitted to you | | `n` | **Rejected** when `n > 1`: 400 `invalid_request`, on `/v1/chat/completions` and `/v1/messages`. The normalizer emits one choice; billing a garbled multi-choice response would be dishonest. The images route takes `n` up to 4 — see [/docs/api-images](/docs/api-images) | | `temperature` | Forwarded when ≤ 1. OpenAI's 0..2 range does not map to Anthropic's 0..1: `temperature > 1` is a typed 400, never a silent clamp | | `top_p` | Forwarded only when `temperature` is absent (Anthropic documents them as mutually exclusive; `temperature` wins, the drop is recorded) | | `parallel_tool_calls: false` | → `tool_choice.disable_parallel_tool_use: true` | | `tools[].function.strict` | Mapped 1:1 | | `response_format`, `reasoning_effort`, `prediction`, `audio`, `modalities` | **Typed 400** naming the param — the gateway does not map these across dialects yet and will not drop them silently: pinning JSON mode or a reasoning budget must not silently produce a different answer | | Trailing `assistant` message | **Typed 400**: Anthropic treats it as an assistant-prefill request, which current Claude models reject | | Everything else (`logprobs`, `seed`, penalties, …) | **Dropped and recorded** — never forwarded | ### Anthropic surface → OpenAI-dialect provider | Parameter | Handling | |---|---| | `model`, `stream`, `temperature`, `top_p` | Copied verbatim | | `max_tokens` | → `max_completion_tokens` (the modern OpenAI field; legacy `max_tokens` is rejected by o-series/GPT-5 reasoning models). Set to 4096 by the gateway when absent, as above | | `system` (string or text blocks) | → one leading `system` message | | `stop_sequences` | → `stop` | | `messages` | Translated: `tool_use` ↔ `tool_calls`, `tool_result` blocks → `role:"tool"` messages (one per result) | | `tools`, `tool_choice` | Translated: `auto`→`"auto"`, `any`→`"required"`, `none`→`"none"`, `{type:"tool",name}`→`{function:{name}}`. `tool_choice.disable_parallel_tool_use: true` → `parallel_tool_calls: false`; `tools[].strict` maps 1:1. Server tools (computer use, web search, bash, text editor) are a **typed 400** — an OpenAI-dialect deployment cannot execute them | | `thinking`, `output_config` | **Typed 400** naming the param (not mappable across dialects yet; native passthrough on Anthropic-dialect routes) | | `top_k`, `metadata` | **Dropped and recorded** (native passthrough on Anthropic-dialect routes). On a house route, `metadata.user_id` is used as the end-user id and is not recorded (see [Sessions and prompt caching](#sessions-and-prompt-caching)) | | `cache_control` (nested and top-level) | To OpenRouter: kept on the OpenAI text parts, and the top-level field too. A mark on a tool definition or a `tool_use` block has no place in OpenAI format: **dropped and recorded**, e.g. `tools[0].cache_control`. To OpenAI direct and Azure: **dropped and recorded with the full path**, e.g. `messages[0].content[0].cache_control` | | `tool_result.is_error` (nested) | **Dropped and recorded with its full path** | > [!NOTE] > "Dropped" means exactly that: the parameter is never forwarded — and every > drop is **recorded**, per request. The full paths appear in the > `x-tm-dropped-params` response header, in each attempt's `dropped_params` in > [`GET /v1/generation`](/docs/api-usage), and in the (content-free) ledger. > Send `"provider": {"require_parameters": true}` to turn any would-be drop > into a typed 400 naming the first dropped parameter instead of a dispatch. > Dropping only ever applies to *parameters*, never content. ### Content is never silently dropped Parameters are droppable because they don't get billed; content is not. A provider that drops a part still bills the request, so two rules apply: - **A part goes only to a model that takes it.** Each listing declares what a message to the model may carry: text, plus any of image, file (a document such as a PDF), audio and video. `GET /v1/models` shows what a model takes for your key in `architecture.input_modalities`. A model's page lists only the inputs that every listing of the model takes, so it can show less than the gateway accepts. The gateway sends a part only to a deployment whose listing takes its kind. A part that no deployment of the model takes is a **typed 400 naming the exact field**, on both surfaces, before anything is reserved or sent: `messages[0].content[1]: model "deepseek/deepseek-v4-flash" does not take image input; it takes text`. The parts checked are `image_url`, `file`, `input_audio` and `video_url` (or `input_video`) on the OpenAI surface, and `image` and `document` blocks on the Anthropic surface, also inside a `tool_result`. A route on your own provider key (BYOK) declares nothing, so the gateway sends every part to it as you sent it. - **A translation names what it cannot carry.** On a translated route, content the translation cannot represent — image parts, audio parts, unknown block types — is a **typed 400 naming the exact field** (`messages[2].content[0]` and so on), never a silent omission. One deliberate exception: assistant `thinking` / `redacted_thinking` blocks echoed back on the Anthropic surface are dropped and recorded rather than rejected. The gateway's own surface emits those blocks (translated from upstream reasoning deltas), so echoing a transcript it produced must not 400 — but thinking content has no OpenAI-dialect equivalent and is never forwarded. ## Sessions and prompt caching A provider keeps a prompt's cache on one machine or host. The gateway passes on what each provider needs to send a conversation back there (decision D23). **Session id.** The gateway takes the first valid value of these, in this order: | Order | Input | Where | |---|---|---| | 1 | `session_id` | Body | | 2 | `x-session-id` | Header | | 3 | `x-session-affinity` | Header | | 4 | `x-claude-code-session-id` | Header | A valid value has 1 to 256 printable ASCII characters. The body field `prompt_cache_key` follows the same rule. An invalid value is ignored, never an error: the gateway uses the next input and records the ignored one as `session_id (invalid)`, `prompt_cache_key (invalid)` or `header: (invalid)`. With `provider.require_parameters: true`, an ignored input is a 400, as for any drop. Valid session inputs are never recorded as dropped. Each route gets its own field: | Route | Field sent | Value | |---|---|---| | OpenAI (house or BYOK), Azure | `prompt_cache_key` | Your `prompt_cache_key`, else the session id. Nothing if you sent neither | | OpenRouter | `session_id` | The session id, else your `prompt_cache_key`, else an id that the gateway makes from your organization, the model and the opening messages. Your `prompt_cache_key` goes too | | Anthropic, Bedrock | Nothing | Anthropic has no session input: its cache matches the prompt prefix | Values go unchanged. A value longer than the provider accepts is replaced by `tm-` followed by 43 base64url characters: a hash of the value. `prompt_cache_retention` (`"in_memory"` or `"24h"`) goes unchanged to OpenAI and Azure; other routes drop and record it. **End-user ids on house routes.** A house route runs on the marketplace's own provider account, which every customer shares. On a house route, the gateway sends exactly one end-user id: a hash of your organization and your `safety_identifier` (else `user`, else, on `/v1/messages`, `metadata.user_id`), or of your organization alone. | House route | Field sent | Not sent | |---|---|---| | `openai` | `safety_identifier` | `user` | | `openrouter` | `user` | `safety_identifier` | | `anthropic` | `metadata.user_id` | `user`, `safety_identifier` | BYOK routes get your values unchanged: `user` and `metadata` as before, and `safety_identifier` on OpenAI and Azure from `/v1/chat/completions`. Anthropic and Bedrock BYOK routes, and any translation from `/v1/messages`, drop and record it. **Cache marks.** The gateway never adds a `cache_control` mark that you did not send. The ledger prices every cache write at the 5-minute rate, so on a house route `ttl: "1h"` is removed from every mark (the cache then keeps 5 minutes) and recorded as `.cache_control.ttl`, for example `system[0].cache_control.ttl`. `ttl: "5m"` passes. BYOK routes keep `ttl: "1h"`: you pay that provider directly. **Staying on the first route.** When the gateway's own quota pool for your first route is full, a follow-up request (one with an assistant or tool message after a user message) waits up to 5 seconds for it, instead of moving to a route with a cold cache. The response then carries `x-tm-affinity-wait-ms` with the milliseconds waited. ## Usage accounting (how tokens are counted) | Concept | OpenAI surface | Anthropic surface | |---|---|---| | Billable input | `prompt_tokens` — cache-inclusive: input + cache reads + cache writes | `input_tokens` — excludes cache reads and writes | | Cache reads | `prompt_tokens_details.cached_tokens` | `cache_read_input_tokens` | | Cache writes | `prompt_tokens_details.cache_write_tokens` | `cache_creation_input_tokens` | | Output | `completion_tokens` | `output_tokens` | | Reasoning | `completion_tokens_details.reasoning_tokens` (a subset of output) | not separately reported | Cross-dialect translation converts exactly by this table; `total_tokens = prompt_tokens + completion_tokens`. - **Usage is always on.** Streams and non-streams both report it. The gateway injects `stream_options.include_usage: true` into every streamed OpenAI-dialect dispatch itself — without it those streams carry no usage at all. The only `stream_options` value that changes what you receive is an explicit `include_usage: false`, which suppresses the usage chunk (you are billed the same). - **`cache_write_tokens` appears on both surfaces.** Cache writes bill at their own rate, so the split is always visible: on billed OpenAI-surface responses `prompt_tokens_details.cache_write_tokens` is always present (0 when the provider reports none), stream and non-stream alike. - **Every billed response carries `usage.cost` (USD)** — in the final usage chunk on streams, in the body otherwise — computed with the exact integer micro-USD math the ledger settles with, and covering the *full* request debit, including attempts that failed over before your answer started. You can always recompute your bill from the wire. - **Rounding favors you.** Prices are micro-USD per million tokens; the pre-flight reservation rounds up, settlement rounds down. > [!WARNING] > One documented limitation: an Anthropic-surface stream that dies mid-answer > has no legal wire slot for usage (`message_delta` is the only one, and it > never arrives). Recompute those from > `GET https://api.routerplus.com/v1/generation?id=REQUEST_ID`, which returns every > physical attempt with tokens and settled cost. The OpenAI surface has no such > gap — known usage and cost are emitted *before* the terminal error event. ## Streams OpenAI-surface event order: role-priming delta (sent as soon as the upstream proves alive) → content deltas → finish chunk → usage chunk → `data: [DONE]`. `[DONE]` is never sent after an error, and nothing ever follows a terminal event. - **Keep-alives.** After output has started, any silence of 15 seconds or more gets a keep-alive frame so proxies and clients don't kill an idle-but-healthy stream (slow reasoning models): an SSE comment (`: processing`) on the OpenAI surface, a native `ping` event on the Anthropic surface. Both are spec-ignorable — SDKs skip them without code changes. - **The commit boundary.** Before any semantic output reaches you, provider failures are handled by invisible zero-backoff failover — you see one response, one monotonic event sequence, and `x-tm-attempts` counts what it took. After output has started, the gateway **never** switches providers: an upstream death becomes one terminal error event inside the 200 stream, and retry semantics stay yours. A refusal is never rerouted to another provider. - **The request deadline.** 450 seconds from dispatch, or a routing policy's `timeout_ms`. A stream still open then gets the same terminal error event. - Mid-stream terminal shape: OpenAI surface — a full chunk envelope with `finish_reason: "error"` and a top-level `error` object; Anthropic surface — an `event: error` frame. ## `finish_reason` and `native_finish_reason` Translated responses map finish reasons both ways: | Anthropic `stop_reason` | OpenAI `finish_reason` | |---|---| | `end_turn` | `stop` | | `stop_sequence` | `stop` | | `max_tokens` | `length` | | `tool_use` | `tool_calls` | | `refusal` | `content_filter` | The mapping is lossy (two stop reasons both become `stop`), so translated responses on the OpenAI surface carry the provider's raw value verbatim in `choices[0].native_finish_reason` — on the non-stream body and on the stream's finish chunk. When it differs from `finish_reason`, the native value is what the provider actually said — trust it for analytics. Passthrough responses are relayed verbatim, so their `finish_reason` is already native. ## Errors Every error body is generated by the gateway in the dialect you called with, and carries a stable `error_type` with the canonical class — `error.metadata.error_type` on the OpenAI surface, `error.error_type` on the Anthropic surface — plus the dialect's native fields so your SDK's built-in handling works untouched. On the Anthropic surface that means real native type strings (`billing_error`, `rate_limit_error`, `authentication_error`, `not_found_error`, `api_error`, `invalid_request_error`), so Anthropic SDK retry classification behaves. A provider's own error body is never relayed; the provider's HTTP status rides in `x-tm-upstream-status`. The body's `metadata` also names where the failure came from (`origin`, and the limit that refused it when one did) — see [Errors](/docs/errors). Canonical classes and how routing treats them: | Upstream signal | `error_type` | Failover | |---|---|---| | 429 (Retry-After honored, capped at 60s) | `rate_limit` | yes | | 401/402/403 from a provider (our account problem, not yours) | `auth` | yes | | Context/length errors | `context_overflow` | no — fix the request or pick a bigger window | | Content policy / refusal | `content_policy` | **never** — a refusal is not rerouted | | 404 / model not found upstream | `model_unavailable` | yes | | 5xx / 529 / overloaded | `upstream_error` | yes | | Other 4xx | `upstream_error` | no | On `POST /v1/decisions` some of these signals read differently, by their status and the provider's own error code, never by the provider's name: Levanto's 402, Fastino's 402 or 403, a System One provider's 401, 402 or 403 on our account, and Cloudflare's 429 with code 3036 (our free daily allocation is used up) are a 503 `model_unavailable` with `metadata.retryable: false`. Cloudflare's 413 with code 5021 is a 413 `context_overflow`; Perplexity's 400 that names the model's maximum context length (on Perplexity Decider v1 27B) is a 400 `context_overflow`, by the context rule above; and a System One provider's 408 is a retryable 504 `upstream_error`. A decisions provider's 429 never counts against the deployment's health circuit. See [POST /v1/decisions](/docs/api-decisions). Out-of-credits is deliberately a 429 `insufficient_quota` (not a 402): byte-compatible with what OpenAI's SDKs already recognize and back off on. Gateway-authored error bodies never echo flagged input content. The full per-status remediation table is on [Errors](/docs/errors). ## Response headers | Header | When | Meaning | |---|---|---| | `x-request-id` | every call | the id `GET /v1/generation?id=` audits | | `x-tm-provider` | served requests | deployment that served the answer | | `x-tm-attempts` | served / failed dispatches | physical dispatches, failovers included | | `x-tm-upstream-status` | when an upstream responded | the provider's own HTTP status | | `x-tm-error-code` | failures | canonical class | | `x-tm-error-origin` | limit, provider and infrastructure failures | where it came from: `gateway_admission`, `upstream_quota`, `upstream`, `gateway_infrastructure` or `authorization` | | `x-tm-limit-scope`, `x-tm-limit-kind`, `x-tm-limit-id` | a refused limit | which limit refused the request — see [Limits and capacity](/docs/admission) | | `x-tm-cap-reset` | spend-cap 429s | exact instant the monthly cap resets (first of next UTC month) | | `x-tm-dropped-params` | when non-empty | comma-joined names of parameters the gateway stripped (D8 §2: drops are recorded, never silent) | | `x-tm-upstream-model` | aggregator swaps | the id the gateway actually sent upstream | | `x-tm-served-by` | aggregator routes that name it | the provider the aggregator used, e.g. `Amazon Bedrock` | | `x-tm-route-plan-id` | billed requests | the id of the route plan that chose the deployments | | `x-tm-admission-mode` | billed requests | `shared`, or `bounded_local` while the shared limit store is unreachable | | `x-tm-remaining-rpm`, `x-tm-remaining-tpm` | admitted requests | the smallest remaining allowance across the limits the request claimed | | `retry-after` | 429/503 | seconds to wait; a provider's value is honored, capped at 60s | Browsers can read every header above except `x-tm-error-code` and `x-tm-route-plan-id`, which are not CORS-exposed; use the body's `error_type` in browser code. Optional, content-free app attribution: send `HTTP-Referer` and `X-Title` to identify your app in per-app analytics. ## Out of scope in v1 Honest errors or recorded drops — nothing silent: - Image, file, audio and video **input** have no cross-dialect translation: typed 400 on translated routes. On a passthrough route a part reaches the provider as sent when the model takes its kind, and is a typed 400 when no deployment of the model takes it; `GET /v1/models` lists the inputs a model takes. - Image and video **generation** are in scope on their own surfaces (D8 §7.6, D17, D18); image edits, variations and streaming, and image-to-video, are not — see [POST /v1/images/generations](/docs/api-images) and [POST /v1/videos](/docs/api-videos). - `n > 1` choices: rejected 400 on the chat surfaces (the images route takes `n` up to 4). - No model id variants — ids containing `:` are not in the catalog and 404 like any unlisted id, and there is never a silent substitute. - On passthrough routes, the forwarded parameters reach the provider as you sent them, and the provider's behavior is the provider's. The guaranteed surface is exactly what this page documents. See also: [Migrate](/docs/migration) · [Quickstart](/docs/quickstart) · [Errors](/docs/errors) --- Source: https://app.routerplus.com/docs/for-providers.md # List your models Sell your inference through the marketplace: one model document, a published conformance bar, pass-through pricing, and automatic payment on the schedule you choose. Two things to know up front: - **Pricing is pass-through.** Buyers pay your listed price; we never mark it up. Your price is your price. - **Routing is earned.** Traffic flows on live health and published metrics — the same numbers you can see — not on negotiation. **To apply:** email **providers@evo-hq.com** with your model document attached. Onboarding is white-glove today: a human plus the automated suite take it from there, usually same-week. ## The model document Everything we need is one JSON document — your endpoint, models, prices, capacity, and data policy. Where OpenRouter polls a `/v1/models` endpoint you host, we accept the same information **pushed** as a versioned manifest: you send updates (new models, price changes) when you choose, and nothing changes under you between updates. Validate against the published JSON Schema before sending: [`/docs/provider-manifest.schema.json`](/docs/provider-manifest.schema.json). ```json { "manifest_version": "1", "provider": { "id": "acme", "name": "Acme Inference", "privacy_policy_url": "https://acme.example/privacy", "terms_of_service_url": "https://acme.example/terms", "status_page_url": "https://status.acme.example", "prompt_logging": "none" }, "endpoint": { "base_url": "https://api.acme.example/v1", "dialect": "openai", "api_key_env": "ACME_API_KEY" }, "settlement": { "mode": "arrears", "interval": "weekly", "billing_contact": "billing@acme.example" }, "models": [ { "id": "acme-large-1", "display_name": "Acme Large 1", "context_length": 128000, "max_output_tokens": 16384, "streaming": true, "is_ready": false, "pricing": [ { "type": "prompt", "unit": "token", "cost_usd_per_million": "0.90" }, { "type": "completion", "unit": "token", "cost_usd_per_million": "3.60" } ], "capacity": [ { "type": "completion", "unit": "token", "per": "minute", "value": 2000000 } ] } ] } ``` ### Identity `id` is the **exact** identifier we send when calling your API — never an alias. `display_name` is what buyers see. A listing is text (chat completions) unless `output_modality` says `"image"`, `"video"` or `"decisions"`. An image listing is served on `POST /v1/images/generations` only; it needs no `context_length` or `streaming`, and must declare `max_output_tokens` (the per-image ceiling the reservation is computed from) plus `supported_parameters` for `n` and a `size` or `aspect_ratio` enum. A video listing is served on `POST /v1/videos`; its `max_output_tokens` is the ceiling per second, and `seconds` must be declared with a default. A decisions listing (a structured decision model that returns typed answers to questions about a state) is served on `POST /v1/decisions`; it needs a `context_length` but no `streaming`, and your endpoint must declare `decisions_path`, the path we POST the request to. If your API picks the model by its URL, each decisions model declares its own `decisions_path` instead (Cloudflare Workers AI: `/clef` and `/clef-flash`), and then the endpoint needs none; the body still carries the model's `upstream_id` (or `id`) as `model`. The model declares `decisions_wire`, the question format your API speaks: `"systemone"` (TypeSafe's System One API, which Jev, Mercury Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B and Perplexity Decider v1 27B use; the default when the field is absent) or `"gliner"` (Fastino's GLiNER schema, sent in an OpenAI chat-completions envelope, so `decisions_path` can be `/chat/completions`). Buyers send questions in that format, and the gateway sends them to you in it. A decisions listing is served only on `POST /v1/decisions`, even when its path is a chat path: chat traffic never reaches it. Two listings of one model id must declare the same wire. `decisions_wire` on a listing that is not `"decisions"` is refused. If your API puts every answer and every error inside a transport envelope of its own, declare it as `endpoint.response_envelope`, and the gateway opens it before it reads the answer. Only a manifest of decisions models may declare it. The one value today is `"cloudflare_v4"`, Cloudflare's REST envelope: a 200 is `{"result": , "success": true, "errors": [], "messages": []}`, and an error is `{"errors": [{"code": ..., "message": ...}], "success": false, "result": {}}`. A 200 whose envelope does not say `success: true` is a failed answer, and nothing is billed for it. Two more fields tell the gateway how your System One API reads a body. They are booleans on a decisions model whose `decisions_wire` is `"systemone"` (or absent), and are refused on any other listing. Absent means `false`. - `decisions_state_per_question`: your API runs one prompt per question, each with the whole state, and bills the state once per question (Perplexity's decider does). The gateway then holds the state once per question when it reserves money and tokens for a request, so a request with many questions is held at what it will cost. - `decisions_image_parts`: your API reads an object whose `type` is `"image_url"`, at any depth of the state or of a question, as an image. A decisions listing takes text and JSON, so the gateway refuses such a body with a typed 400 before any money is reserved, and it never reaches you. A chat listing may declare `input_modalities`: what a message to the model may carry. The list always has `"text"`, plus any of `"image"`, `"file"` (a document such as a PDF), `"audio"` and `"video"`. Buyers see it on the model page and filter the catalog by it, and the API lists it in `/v1/models`. The model page shows an input only when every listing of that model declares it. The gateway sends a part only to a listing that declares its kind: a part your listing does not declare never reaches you, and a part that no listing of the model declares is refused with a typed 400. A declared part goes to you as sent when the buyer calls the surface that speaks your dialect; across dialects the gateway refuses it with a typed 400 and does not translate it. Absent means text only. An image, video or decisions listing takes a text prompt, and the field is refused on it. ### Serving a model the marketplace already lists If you serve a model another provider also lists (the marketplace id `claude-sonnet-4-5`, say), list it under **that** id and set `upstream_id` to the name your endpoint expects: ```json { "id": "claude-sonnet-4-5", "upstream_id": "anthropic/claude-sonnet-4.5", "...": "..." } ``` Requests to you carry `upstream_id`; conformance check C7 expects you to echo it. Prices must be **identical** to the other providers of that id — the catalog build rejects a mismatch (pass-through pricing means one price per model). `provider.priority` orders providers of one model: `0` for a first-party lab, `1` for an aggregator fallback; routing tries lower first. ### Endpoint `dialect` is `"openai"` (Chat Completions) or `"anthropic"` (Messages) — one is enough. Buyers reach your models from **both** marketplace surfaces regardless; the gateway translates. `api_key_env` names the secret slot for the key you issue us: it never appears in a manifest, page, or log. ### Pricing Entries are `{type, unit, cost_usd_per_million}` with costs as decimal **strings** in USD (never floats). Types: `prompt`, `cached_prompt`, `cache_write`, `completion`, `internal_reasoning`; unit is `token`. - **Declare what you charge.** A SKU you omit bills those tokens at your base `prompt`/`completion` rate — a missing SKU never means free, and never means a surprise for either side. - **Price changes are append-only.** A new effective-dated row, never a rewrite. Every buyer request bills at the snapshot in force when it was dispatched, so your statement and their bill can never disagree about history. - Flat per-token pricing only for now: conditional overrides, time-of-day windows, and per-request SKUs are not yet supported (see the OpenRouter mapping below). ### Capacity Entries are `{type, unit, per, value}` — `request`, `prompt` or `completion` limits per `minute`/`hour`/`day` window (`unit: "token"` is required for the two token types). Identity is type + window; duplicates are rejected at validation. Absent capacity means undeclared, not zero. Capacity is information for us, not a setting. Validation checks it, but routing and admission do not read it. The limits that admission enforces for your deployment are a house pool that our operators set by hand from what you declare and what you tell us (see [Limits and capacity](/docs/admission)). Above that pool, your own 429 is the backstop: it pauses the pool for its `retry-after` (1 to 60 seconds). ### Parameters `supported_parameters` may be declared per model, and it is **enforced on the image and video routes**: a value outside a declared enum is a 400 naming the field. An image listing must declare `n` and a `size` or `aspect_ratio` enum, or it fails validation; `quality` and `resolution` are optional but must be enums when declared. On chat routes the gateway applies its own per-dialect allowlists, and **every parameter it strips is recorded per-request** (visible to the buyer in the `x-tm-dropped-params` header and their audit endpoint). Declaring accurately now means chat enforcement picks it up when it lands. ### Data policy `prompt_logging` is `"none"` or `"retained"`, published verbatim on your model pages next to your privacy policy and terms URLs. Say the true thing — buyers filter on it. Our own side is content-free by design: the marketplace never stores prompts or responses, so your data policy is the only one in the path. ### Lifecycle - **`is_ready: false`** stages a model: validated, conformance-testable, listed nowhere, routed never. Flip it when you are ready — a model proves itself before going live, not after. - **`deprecation_date`** (ISO date): past it, the model leaves the catalog automatically. - **`released`** (ISO date): the day the model came out. The Models page and the playground list models newest first after a few featured ones, so set it when you add a model: the lab's release date, or the day the model first listed publicly. - Date-pinned requests (`acme-large-1-20260901`) resolve to your listed model for routing and billing, and the buyer's original id reaches your API verbatim — you decide whether to serve the snapshot. ## The conformance bar Where OpenRouter runs unpublished "baseline tests," our bar is **published** — seven checks, run against every model on both the stream and non-stream paths, at listing time and on every catalog change. We share the report with you. - **C1** — a non-stream completion returns provider-reported usage tokens. Usage is the billable record; no usage, no listing. - **C2** — streamed responses are valid SSE, every frame parseable. - **C3** — streams end with their dialect's terminal marker (`[DONE]` / `message_stop`). - **C4** — usage tokens are reported in-stream. - **C5** — an unknown model returns a parseable 4xx JSON error, not a 200 and not HTML. - **C6** — a billed response actually contains assistant content. Usage counters without an answer is the one failure we will never pass. - **C7** — the response reports the model that was asked for. Resolving an undated id to your own dated snapshot is fine; any other equivalence must be declared in `resolves_to`. Silently substituting a different model is disqualifying. An **image** listing (`output_modality: "image"`) is non-stream only, so C2–C4 do not apply; it runs C5 plus three image checks, each of which buys one real render at the model's top declared quality: - **I1** — a generation returns decodable base64 and provider-reported usage tokens. - **I2** — the tokens billed for that render fit the declared per-image ceiling (`max_output_tokens`). An under-declared ceiling is a house loss on every request. - **I3** — an undeclared size returns a parseable 4xx JSON error. - **I4** (opt-in, `--image-ceiling`) — I2 at every declared size and at `size: auto`. A **video** listing (`output_modality: "video"`) runs C5 plus three video checks: - **V1** — a video job at the cheapest tier finishes, reports its cost and downloads as an MP4. - **V2** — the declared per-second ceiling covers the reported cost. - **V3** — an undeclared shape is refused, not silently substituted. A **decisions** listing (`output_modality: "decisions"`) runs two checks: - **D1** — one call with three questions comes back as typed answers, with usage. On the `systemone` wire: one Choice, one Score and one Noul question, and `usage.input_tokens`. On the `gliner` wire: one yes/no task and one three-way task in one schema (sent with `store: false`), the message content parsing to one JSON object with a labelled answer under each task name, and `prompt_tokens` in the usage. D1 calls each model at its own `decisions_path` when the model declares one, else at the endpoint's. With `endpoint.response_envelope: "cloudflare_v4"`, the 200 must say `success: true`, and D1 reads the answer and the usage inside `result`. When D1 gets another status, the report shows the first 300 characters of your body. - **C5** — an unknown model yields a parseable 4xx error. FastAPI's `{"detail": ...}` envelope counts as a described error, and so does Cloudflare's `{"errors": [{"code": ..., "message": ...}]}` (or one such object in place of the list) when an entry has a message. For every model, the report also lists any `Ratelimit`, `Ratelimit-Policy` or `Retry-After` header your endpoint sent, or says that it sent none. These headers are reported, never judged. A provider that reports its own cost per render instead of tokens declares `endpoint.billing: "reported_cost"`; the ledger then charges exactly the cost reported, and the image and video checks judge that cost against the ceiling. ## Wire requirements - HTTPS endpoint speaking OpenAI Chat Completions **or** Anthropic Messages, with streaming. - Provider-reported usage tokens on both paths (C1/C4). - Honest errors: return early 429s at capacity — queueing wrecks your published latency and your buyers' experience; 4xx JSON for bad requests. - **Streaming:** response headers within 20 seconds or the attempt fails over; send SSE keep-alive comments through long pre-first-byte gaps, and stream tokens as soon as they exist. Non-stream requests are exempt — they get a long generation budget instead. - After first output we never switch providers mid-answer — a stall is surfaced to the buyer as yours. Stream steadily. ## How uptime is counted Exact accounting, same numbers routing uses: ```text uptime = completed ÷ (completed + counted failures), trailing 24h ``` - **Counted against you:** provider 401/402/403, all 5xx, mid-stream deaths, incomplete streams. - **Never counted:** buyer-caused requests — 400 and 413 (bad input), 429 (their rate limit), and client cancellations mid-stream. - **Published only after 100+ requests.** Below that the pages say `n<100` — we never print a fake 100%, yours or anyone's. - Failures open a per-deployment **health circuit**: cooldown, then exactly one half-open probe; recovery closes it. While everything for a model cools down we still send a last-resort attempt rather than strand buyers. Content-policy refusals are never retried and never failed over. ## Performance We publish **TTFB p50/p95** (from merged t-digests) per model — the same figures the router sees. Two behaviors dominate them: return early 429s instead of queueing, and start streaming immediately. ## How you get paid Pick a settlement mode in your manifest — the equivalent of OpenRouter's "auto top-up or invoicing" requirement, both directions first-class: - **`prepaid`** — we deposit with you ahead of traffic; usage draws it down, and we monitor balance against burn so routing never hits a dry account. - **`arrears`** — we route first and pay on your interval: `half_daily`, `daily`, `weekly`, or `monthly`. Your statement is generated from the **same append-only ledger that bills buyers** — token counts and USD per period, auditable to the individual attempt, with per-attempt byte counts recorded as a tokenizer-independent bound for disputes. ## What you get - Buyer traffic from both wire surfaces the day your models go live. - A public model page per model: prices, context, your data-policy flag, live honest metrics. - Per-period settlement statements that reconcile to the token. - Degraded providers get circuits and time to recover, not silent delisting. ## Coming from OpenRouter? The concepts map directly; here is the translation: | OpenRouter | Here | |---|---| | Polled `/v1/models` model documents | Pushed manifest (this page) — same information, updated when you choose | | `is_ready`, `deprecation_date` | Same names, same semantics | | Baseline tests (unpublished) | C1–C7, I1–I4 and V1–V3, published above, report shared with you | | Uptime after 100+ requests, user errors excluded | Same — formula published above | | Data-policy / may-train disclosure | `prompt_logging` + your policy URLs, shown to buyers | | Auto top-up or invoicing | `prepaid` or `arrears` with ledger-audited statements | | Conditional pricing, time-of-day windows, `discount_to_user`, `:free` variants | Not yet — flat per-token pricing only | | Multimodal listings | Image and video generation, and decision models, via `output_modality`; image, file, audio and video inputs to chat models via `input_modalities` (passed through in your own dialect, never translated; a kind you do not declare never reaches you) | | Provider dashboard | Not yet — statements and conformance reports by email while the console is built | ## Apply Email **providers@evo-hq.com** with your manifest attached (validate it against [the schema](/docs/provider-manifest.schema.json) first). We run the suite against your endpoint, share the report, stage your models, and flip them live together.