Wire compatibility
The D8 conventions your bill is computed from.
This page is the published form of decision D8 — the conventions your bill is computed from. Changing anything here is treated internally as a billing-contract change and requires a recorded decision.
Two surfaces, one catalog
| Surface | Endpoint | Auth |
|---|---|---|
| OpenAI-compatible | POST https://api.routerplus.com/v1/chat/completions | Authorization: Bearer or x-api-key |
| Anthropic-compatible | POST https://api.routerplus.com/v1/messages | Authorization: Bearer or x-api-key |
| OpenAI Images | POST https://api.routerplus.com/v1/images/generations | same |
| OpenAI Videos | POST https://api.routerplus.com/v1/videos | same |
Every cataloged chat model is callable from both chat surfaces; image models answer on the Images surface only, video models on the Videos surface only. GET https://api.routerplus.com/v1/models lists what your key can serve — OpenAI list shape by default, Anthropic shape when you send an anthropic-version header.
One distinction drives everything below:
- Passthrough route — your dialect matches the serving provider's (OpenAI format → OpenAI-dialect provider, Anthropic format → Anthropic-dialect provider). The parameters your dialect defines are forwarded as you sent them, and the gateway adds
stream_options.include_usageon streamed OpenAI-dialect dispatches. Whatever the provider accepts for those parameters, you get — including native features like Anthropic prompt caching (cache_control),thinking,top_k, and theanthropic-betavalues listed below. A top-level key outside your dialect's parameter set is dropped and recorded, never forwarded. Session ids and cache marks follow Sessions and prompt caching. - Translated route — dialects differ, and the gateway translates the request, the stream, and errors. Translation covers a pinned parameter set; everything else follows the rules below.
A "typed 400" in the translation tables is decided per deployment. When a translation refuses your request but another deployment of the model speaks your own dialect, the request goes there instead; you get the 400 only when no deployment can serve it. Content a model does not take follows the same rule: a deployment that does not take a part is skipped, and the 400 comes only when no deployment takes it (see Content is never silently dropped).
Parameter handling on passthrough routes
The forwarded set per dialect. Anything else at the top level is a recorded drop (see the note under the translation tables).
| Dialect | Forwarded as sent |
|---|---|
| OpenAI → OpenAI-dialect provider | model, messages, stream, stream_options, temperature, top_p, max_tokens, max_completion_tokens, stop, n, tools, tool_choice, parallel_tool_calls, response_format, reasoning_effort, seed, user, logprobs, top_logprobs, frequency_penalty, presence_penalty, logit_bias, metadata, store, prediction |
| Anthropic → Anthropic-dialect provider | model, messages, system, max_tokens, stream, temperature, top_p, top_k, stop_sequences, tools, tool_choice, metadata, thinking, output_config, cache_control |
Some OpenAI-format keys go only to the providers that define them:
prompt_cache_retentionandsafety_identifiergo to OpenAI and Azure.- The top-level
cache_control(automatic caching) goes to OpenRouter. - A
cache_controlmark on a content part passes unchanged to every OpenAI-dialect provider. OpenRouter uses it; OpenAI and Azure accept it and ignore it.
On Bedrock, the top-level cache_control becomes a mark on the last block that can carry one, because Bedrock's API does not take the top-level field. On every route, a top-level cache_control that Anthropic would refuse is dropped and recorded as cache_control: a field other than {"type": "ephemeral"} with an optional ttl of "5m" or "1h", a fifth mark, a 1-hour field after a 5-minute mark, or a field whose TTL differs from the mark on the last block.
Two Anthropic parameters are refused on every route, because stripping them would change who runs what: mcp_servers and container are a typed 400.
The anthropic-beta header is forwarded to Anthropic-dialect providers only for values that do not change how a token is billed: prompt-caching, token-efficient-tools, fine-grained-tool-streaming, interleaved-thinking, claude-code, oauth and computer-use prefixes. Any other value is dropped and recorded as header:anthropic-beta:<value>.
Parameter handling on translated routes
OpenAI surface → Anthropic-dialect provider
| Parameter | Handling |
|---|---|
model, messages, stream, temperature, top_p | Translated / copied verbatim |
max_tokens, max_completion_tokens | Translated. The gateway sets max_tokens: 4096 on any request that carries neither, before dispatch; see POST /v1/chat/completions for the 1–32,768 bound |
stop (string or array) | → stop_sequences |
system / developer messages | → the Anthropic system string, joined in order with blank lines. If a part has a cache_control mark, system becomes text blocks with the same text, and each mark ends a block |
cache_control on a text part | Kept on the Anthropic block. A role:"tool" message's mark goes on its tool_result |
| Other keys of a text part | Not forwarded, and not recorded |
cache_control (top-level) | → the Anthropic top-level field (automatic caching). On Bedrock it becomes a mark on the last block |
tools, tool_choice | Translated: auto→{type:"auto"}, required→{type:"any"}, {function:{name}}→{type:"tool",name}, none→{type:"none"}. With none on a tool-free history, tools are omitted entirely so you don't pay for definitions you forbade using. A tool without parameters gets an empty object schema |
role:"tool" messages | → tool_result blocks; consecutive results merge into one Anthropic user message |
An assistant message with content: null and no tool_calls | Dropped and recorded (OpenAI's own refusal shape); it has nothing Anthropic can carry |
stream_options | Usage is on by default; an explicit stream_options.include_usage: false is honored — you are still billed, but no usage chunk is emitted to you |
n | Rejected when n > 1: 400 invalid_request, on /v1/chat/completions and /v1/messages. The normalizer emits one choice; billing a garbled multi-choice response would be dishonest. The images route takes n up to 4 — see /docs/api-images |
temperature | Forwarded when ≤ 1. OpenAI's 0..2 range does not map to Anthropic's 0..1: temperature > 1 is a typed 400, never a silent clamp |
top_p | Forwarded only when temperature is absent (Anthropic documents them as mutually exclusive; temperature wins, the drop is recorded) |
parallel_tool_calls: false | → tool_choice.disable_parallel_tool_use: true |
tools[].function.strict | Mapped 1:1 |
response_format, reasoning_effort, prediction, audio, modalities | Typed 400 naming the param — the gateway does not map these across dialects yet and will not drop them silently: pinning JSON mode or a reasoning budget must not silently produce a different answer |
Trailing assistant message | Typed 400: Anthropic treats it as an assistant-prefill request, which current Claude models reject |
Everything else (logprobs, seed, penalties, …) | Dropped and recorded — never forwarded |
Anthropic surface → OpenAI-dialect provider
| Parameter | Handling |
|---|---|
model, stream, temperature, top_p | Copied verbatim |
max_tokens | → max_completion_tokens (the modern OpenAI field; legacy max_tokens is rejected by o-series/GPT-5 reasoning models). Set to 4096 by the gateway when absent, as above |
system (string or text blocks) | → one leading system message |
stop_sequences | → stop |
messages | Translated: tool_use ↔ tool_calls, tool_result blocks → role:"tool" messages (one per result) |
tools, tool_choice | Translated: auto→"auto", any→"required", none→"none", {type:"tool",name}→{function:{name}}. tool_choice.disable_parallel_tool_use: true → parallel_tool_calls: false; tools[].strict maps 1:1. Server tools (computer use, web search, bash, text editor) are a typed 400 — an OpenAI-dialect deployment cannot execute them |
thinking, output_config | Typed 400 naming the param (not mappable across dialects yet; native passthrough on Anthropic-dialect routes) |
top_k, metadata | Dropped and recorded (native passthrough on Anthropic-dialect routes). On a house route, metadata.user_id is used as the end-user id and is not recorded (see Sessions and prompt caching) |
cache_control (nested and top-level) | To OpenRouter: kept on the OpenAI text parts, and the top-level field too. A mark on a tool definition or a tool_use block has no place in OpenAI format: dropped and recorded, e.g. tools[0].cache_control. To OpenAI direct and Azure: dropped and recorded with the full path, e.g. messages[0].content[0].cache_control |
tool_result.is_error (nested) | Dropped and recorded with its full path |
"Dropped" means exactly that: the parameter is never forwarded — and every drop is recorded, per request. The full paths appear in the x-tm-dropped-params response header, in each attempt's dropped_params in GET /v1/generation, and in the (content-free) ledger. Send "provider": {"require_parameters": true} to turn any would-be drop into a typed 400 naming the first dropped parameter instead of a dispatch. Dropping only ever applies to parameters, never content.
Content is never silently dropped
Parameters are droppable because they don't get billed; content is not. A provider that drops a part still bills the request, so two rules apply:
- A part goes only to a model that takes it. Each listing declares what a message to the model may carry: text, plus any of image, file (a document such as a PDF), audio and video.
GET /v1/modelsshows what a model takes for your key inarchitecture.input_modalities. A model's page lists only the inputs that every listing of the model takes, so it can show less than the gateway accepts. The gateway sends a part only to a deployment whose listing takes its kind. A part that no deployment of the model takes is a typed 400 naming the exact field, on both surfaces, before anything is reserved or sent:messages[0].content[1]: model "deepseek/deepseek-v4-flash" does not take image input; it takes text. The parts checked areimage_url,file,input_audioandvideo_url(orinput_video) on the OpenAI surface, andimageanddocumentblocks on the Anthropic surface, also inside atool_result. A route on your own provider key (BYOK) declares nothing, so the gateway sends every part to it as you sent it. - A translation names what it cannot carry. On a translated route, content the translation cannot represent — image parts, audio parts, unknown block types — is a typed 400 naming the exact field (
messages[2].content[0]and so on), never a silent omission.
One deliberate exception: assistant thinking / redacted_thinking blocks echoed back on the Anthropic surface are dropped and recorded rather than rejected. The gateway's own surface emits those blocks (translated from upstream reasoning deltas), so echoing a transcript it produced must not 400 — but thinking content has no OpenAI-dialect equivalent and is never forwarded.
Sessions and prompt caching
A provider keeps a prompt's cache on one machine or host. The gateway passes on what each provider needs to send a conversation back there (decision D23).
Session id. The gateway takes the first valid value of these, in this order:
| Order | Input | Where |
|---|---|---|
| 1 | session_id | Body |
| 2 | x-session-id | Header |
| 3 | x-session-affinity | Header |
| 4 | x-claude-code-session-id | Header |
A valid value has 1 to 256 printable ASCII characters. The body field prompt_cache_key follows the same rule. An invalid value is ignored, never an error: the gateway uses the next input and records the ignored one as session_id (invalid), prompt_cache_key (invalid) or header:<name> (invalid). With provider.require_parameters: true, an ignored input is a 400, as for any drop. Valid session inputs are never recorded as dropped.
Each route gets its own field:
| Route | Field sent | Value |
|---|---|---|
| OpenAI (house or BYOK), Azure | prompt_cache_key | Your prompt_cache_key, else the session id. Nothing if you sent neither |
| OpenRouter | session_id | The session id, else your prompt_cache_key, else an id that the gateway makes from your organization, the model and the opening messages. Your prompt_cache_key goes too |
| Anthropic, Bedrock | Nothing | Anthropic has no session input: its cache matches the prompt prefix |
Values go unchanged. A value longer than the provider accepts is replaced by tm- followed by 43 base64url characters: a hash of the value. prompt_cache_retention ("in_memory" or "24h") goes unchanged to OpenAI and Azure; other routes drop and record it.
End-user ids on house routes. A house route runs on the marketplace's own provider account, which every customer shares. On a house route, the gateway sends exactly one end-user id: a hash of your organization and your safety_identifier (else user, else, on /v1/messages, metadata.user_id), or of your organization alone.
| House route | Field sent | Not sent |
|---|---|---|
openai | safety_identifier | user |
openrouter | user | safety_identifier |
anthropic | metadata.user_id | user, safety_identifier |
BYOK routes get your values unchanged: user and metadata as before, and safety_identifier on OpenAI and Azure from /v1/chat/completions. Anthropic and Bedrock BYOK routes, and any translation from /v1/messages, drop and record it.
Cache marks. The gateway never adds a cache_control mark that you did not send. The ledger prices every cache write at the 5-minute rate, so on a house route ttl: "1h" is removed from every mark (the cache then keeps 5 minutes) and recorded as <path>.cache_control.ttl, for example system[0].cache_control.ttl. ttl: "5m" passes. BYOK routes keep ttl: "1h": you pay that provider directly.
Staying on the first route. When the gateway's own quota pool for your first route is full, a follow-up request (one with an assistant or tool message after a user message) waits up to 5 seconds for it, instead of moving to a route with a cold cache. The response then carries x-tm-affinity-wait-ms with the milliseconds waited.
Usage accounting (how tokens are counted)
| Concept | OpenAI surface | Anthropic surface |
|---|---|---|
| Billable input | prompt_tokens — cache-inclusive: input + cache reads + cache writes | input_tokens — excludes cache reads and writes |
| Cache reads | prompt_tokens_details.cached_tokens | cache_read_input_tokens |
| Cache writes | prompt_tokens_details.cache_write_tokens | cache_creation_input_tokens |
| Output | completion_tokens | output_tokens |
| Reasoning | completion_tokens_details.reasoning_tokens (a subset of output) | not separately reported |
Cross-dialect translation converts exactly by this table; total_tokens = prompt_tokens + completion_tokens.
- Usage is always on. Streams and non-streams both report it. The gateway injects
stream_options.include_usage: trueinto every streamed OpenAI-dialect dispatch itself — without it those streams carry no usage at all. The onlystream_optionsvalue that changes what you receive is an explicitinclude_usage: false, which suppresses the usage chunk (you are billed the same). cache_write_tokensappears on both surfaces. Cache writes bill at their own rate, so the split is always visible: on billed OpenAI-surface responsesprompt_tokens_details.cache_write_tokensis always present (0 when the provider reports none), stream and non-stream alike.- Every billed response carries
usage.cost(USD) — in the final usage chunk on streams, in the body otherwise — computed with the exact integer micro-USD math the ledger settles with, and covering the full request debit, including attempts that failed over before your answer started. You can always recompute your bill from the wire. - Rounding favors you. Prices are micro-USD per million tokens; the pre-flight reservation rounds up, settlement rounds down.
One documented limitation: an Anthropic-surface stream that dies mid-answer has no legal wire slot for usage (message_delta is the only one, and it never arrives). Recompute those from GET https://api.routerplus.com/v1/generation?id=REQUEST_ID, which returns every physical attempt with tokens and settled cost. The OpenAI surface has no such gap — known usage and cost are emitted before the terminal error event.
Streams
OpenAI-surface event order: role-priming delta (sent as soon as the upstream proves alive) → content deltas → finish chunk → usage chunk → data: [DONE]. [DONE] is never sent after an error, and nothing ever follows a terminal event.
- Keep-alives. After output has started, any silence of 15 seconds or more gets a keep-alive frame so proxies and clients don't kill an idle-but-healthy stream (slow reasoning models): an SSE comment (
: processing) on the OpenAI surface, a nativepingevent on the Anthropic surface. Both are spec-ignorable — SDKs skip them without code changes. - The commit boundary. Before any semantic output reaches you, provider failures are handled by invisible zero-backoff failover — you see one response, one monotonic event sequence, and
x-tm-attemptscounts what it took. After output has started, the gateway never switches providers: an upstream death becomes one terminal error event inside the 200 stream, and retry semantics stay yours. A refusal is never rerouted to another provider. - The request deadline. 450 seconds from dispatch, or a routing policy's
timeout_ms. A stream still open then gets the same terminal error event. - Mid-stream terminal shape: OpenAI surface — a full chunk envelope with
finish_reason: "error"and a top-levelerrorobject; Anthropic surface — anevent: errorframe.
finish_reason and native_finish_reason
Translated responses map finish reasons both ways:
Anthropic stop_reason | OpenAI finish_reason |
|---|---|
end_turn | stop |
stop_sequence | stop |
max_tokens | length |
tool_use | tool_calls |
refusal | content_filter |
The mapping is lossy (two stop reasons both become stop), so translated responses on the OpenAI surface carry the provider's raw value verbatim in choices[0].native_finish_reason — on the non-stream body and on the stream's finish chunk. When it differs from finish_reason, the native value is what the provider actually said — trust it for analytics. Passthrough responses are relayed verbatim, so their finish_reason is already native.
Errors
Every error body is generated by the gateway in the dialect you called with, and carries a stable error_type with the canonical class — error.metadata.error_type on the OpenAI surface, error.error_type on the Anthropic surface — plus the dialect's native fields so your SDK's built-in handling works untouched. On the Anthropic surface that means real native type strings (billing_error, rate_limit_error, authentication_error, not_found_error, api_error, invalid_request_error), so Anthropic SDK retry classification behaves. A provider's own error body is never relayed; the provider's HTTP status rides in x-tm-upstream-status. The body's metadata also names where the failure came from (origin, and the limit that refused it when one did) — see Errors.
Canonical classes and how routing treats them:
| Upstream signal | error_type | Failover |
|---|---|---|
| 429 (Retry-After honored, capped at 60s) | rate_limit | yes |
| 401/402/403 from a provider (our account problem, not yours) | auth | yes |
| Context/length errors | context_overflow | no — fix the request or pick a bigger window |
| Content policy / refusal | content_policy | never — a refusal is not rerouted |
| 404 / model not found upstream | model_unavailable | yes |
| 5xx / 529 / overloaded | upstream_error | yes |
| Other 4xx | upstream_error | no |
On POST /v1/decisions some of these signals read differently, by their status and the provider's own error code, never by the provider's name: Levanto's 402, Fastino's 402 or 403, a System One provider's 401, 402 or 403 on our account, and Cloudflare's 429 with code 3036 (our free daily allocation is used up) are a 503 model_unavailable with metadata.retryable: false. Cloudflare's 413 with code 5021 is a 413 context_overflow; Perplexity's 400 that names the model's maximum context length (on Perplexity Decider v1 27B) is a 400 context_overflow, by the context rule above; and a System One provider's 408 is a retryable 504 upstream_error. A decisions provider's 429 never counts against the deployment's health circuit. See POST /v1/decisions.
Out-of-credits is deliberately a 429 insufficient_quota (not a 402): byte-compatible with what OpenAI's SDKs already recognize and back off on. Gateway-authored error bodies never echo flagged input content. The full per-status remediation table is on Errors.
Response headers
| Header | When | Meaning |
|---|---|---|
x-request-id | every call | the id GET /v1/generation?id= audits |
x-tm-provider | served requests | deployment that served the answer |
x-tm-attempts | served / failed dispatches | physical dispatches, failovers included |
x-tm-upstream-status | when an upstream responded | the provider's own HTTP status |
x-tm-error-code | failures | canonical class |
x-tm-error-origin | limit, provider and infrastructure failures | where it came from: gateway_admission, upstream_quota, upstream, gateway_infrastructure or authorization |
x-tm-limit-scope, x-tm-limit-kind, x-tm-limit-id | a refused limit | which limit refused the request — see Limits and capacity |
x-tm-cap-reset | spend-cap 429s | exact instant the monthly cap resets (first of next UTC month) |
x-tm-dropped-params | when non-empty | comma-joined names of parameters the gateway stripped (D8 §2: drops are recorded, never silent) |
x-tm-upstream-model | aggregator swaps | the id the gateway actually sent upstream |
x-tm-served-by | aggregator routes that name it | the provider the aggregator used, e.g. Amazon Bedrock |
x-tm-route-plan-id | billed requests | the id of the route plan that chose the deployments |
x-tm-admission-mode | billed requests | shared, or bounded_local while the shared limit store is unreachable |
x-tm-remaining-rpm, x-tm-remaining-tpm | admitted requests | the smallest remaining allowance across the limits the request claimed |
retry-after | 429/503 | seconds to wait; a provider's value is honored, capped at 60s |
Browsers can read every header above except x-tm-error-code and x-tm-route-plan-id, which are not CORS-exposed; use the body's error_type in browser code.
Optional, content-free app attribution: send HTTP-Referer and X-Title to identify your app in per-app analytics.
Out of scope in v1
Honest errors or recorded drops — nothing silent:
- Image, file, audio and video input have no cross-dialect translation: typed 400 on translated routes. On a passthrough route a part reaches the provider as sent when the model takes its kind, and is a typed 400 when no deployment of the model takes it;
GET /v1/modelslists the inputs a model takes. - Image and video generation are in scope on their own surfaces (D8 §7.6, D17, D18); image edits, variations and streaming, and image-to-video, are not — see POST /v1/images/generations and POST /v1/videos.
n > 1choices: rejected 400 on the chat surfaces (the images route takesnup to 4).- No model id variants — ids containing
:are not in the catalog and 404 like any unlisted id, and there is never a silent substitute. - On passthrough routes, the forwarded parameters reach the provider as you sent them, and the provider's behavior is the provider's. The guaranteed surface is exactly what this page documents.
See also: Migrate · Quickstart · Errors
Markdown source for agents: /docs/compat.md · index at /llms.txt