Console
Guides/Wire compatibility

Wire compatibility

The D8 conventions your bill is computed from.

/llms.txt

This page is the published form of decision D8 — the conventions your bill is computed from. Changing anything here is treated internally as a billing-contract change and requires a recorded decision.

Two surfaces, one catalog

SurfaceEndpointAuth
OpenAI-compatiblePOST https://api.routerplus.com/v1/chat/completionsAuthorization: Bearer or x-api-key
Anthropic-compatiblePOST https://api.routerplus.com/v1/messagesAuthorization: Bearer or x-api-key
OpenAI ImagesPOST https://api.routerplus.com/v1/images/generationssame
OpenAI VideosPOST https://api.routerplus.com/v1/videossame

Every cataloged chat model is callable from both chat surfaces; image models answer on the Images surface only, video models on the Videos surface only. GET https://api.routerplus.com/v1/models lists what your key can serve — OpenAI list shape by default, Anthropic shape when you send an anthropic-version header.

One distinction drives everything below:

  • Passthrough route — your dialect matches the serving provider's (OpenAI format → OpenAI-dialect provider, Anthropic format → Anthropic-dialect provider). The parameters your dialect defines are forwarded as you sent them, and the gateway adds stream_options.include_usage on streamed OpenAI-dialect dispatches. Whatever the provider accepts for those parameters, you get — including native features like Anthropic prompt caching (cache_control), thinking, top_k, and the anthropic-beta values listed below. A top-level key outside your dialect's parameter set is dropped and recorded, never forwarded. Session ids and cache marks follow Sessions and prompt caching.
  • Translated route — dialects differ, and the gateway translates the request, the stream, and errors. Translation covers a pinned parameter set; everything else follows the rules below.

A "typed 400" in the translation tables is decided per deployment. When a translation refuses your request but another deployment of the model speaks your own dialect, the request goes there instead; you get the 400 only when no deployment can serve it. Content a model does not take follows the same rule: a deployment that does not take a part is skipped, and the 400 comes only when no deployment takes it (see Content is never silently dropped).

Parameter handling on passthrough routes

The forwarded set per dialect. Anything else at the top level is a recorded drop (see the note under the translation tables).

DialectForwarded as sent
OpenAI → OpenAI-dialect providermodel, messages, stream, stream_options, temperature, top_p, max_tokens, max_completion_tokens, stop, n, tools, tool_choice, parallel_tool_calls, response_format, reasoning_effort, seed, user, logprobs, top_logprobs, frequency_penalty, presence_penalty, logit_bias, metadata, store, prediction
Anthropic → Anthropic-dialect providermodel, messages, system, max_tokens, stream, temperature, top_p, top_k, stop_sequences, tools, tool_choice, metadata, thinking, output_config, cache_control

Some OpenAI-format keys go only to the providers that define them:

  • prompt_cache_retention and safety_identifier go to OpenAI and Azure.
  • The top-level cache_control (automatic caching) goes to OpenRouter.
  • A cache_control mark on a content part passes unchanged to every OpenAI-dialect provider. OpenRouter uses it; OpenAI and Azure accept it and ignore it.

On Bedrock, the top-level cache_control becomes a mark on the last block that can carry one, because Bedrock's API does not take the top-level field. On every route, a top-level cache_control that Anthropic would refuse is dropped and recorded as cache_control: a field other than {"type": "ephemeral"} with an optional ttl of "5m" or "1h", a fifth mark, a 1-hour field after a 5-minute mark, or a field whose TTL differs from the mark on the last block.

Two Anthropic parameters are refused on every route, because stripping them would change who runs what: mcp_servers and container are a typed 400.

The anthropic-beta header is forwarded to Anthropic-dialect providers only for values that do not change how a token is billed: prompt-caching, token-efficient-tools, fine-grained-tool-streaming, interleaved-thinking, claude-code, oauth and computer-use prefixes. Any other value is dropped and recorded as header:anthropic-beta:<value>.

Parameter handling on translated routes

OpenAI surface → Anthropic-dialect provider

ParameterHandling
model, messages, stream, temperature, top_pTranslated / copied verbatim
max_tokens, max_completion_tokensTranslated. The gateway sets max_tokens: 4096 on any request that carries neither, before dispatch; see POST /v1/chat/completions for the 1–32,768 bound
stop (string or array)→ stop_sequences
system / developer messages→ the Anthropic system string, joined in order with blank lines. If a part has a cache_control mark, system becomes text blocks with the same text, and each mark ends a block
cache_control on a text partKept on the Anthropic block. A role:"tool" message's mark goes on its tool_result
Other keys of a text partNot forwarded, and not recorded
cache_control (top-level)→ the Anthropic top-level field (automatic caching). On Bedrock it becomes a mark on the last block
tools, tool_choiceTranslated: auto→{type:"auto"}, required→{type:"any"}, {function:{name}}→{type:"tool",name}, none→{type:"none"}. With none on a tool-free history, tools are omitted entirely so you don't pay for definitions you forbade using. A tool without parameters gets an empty object schema
role:"tool" messages→ tool_result blocks; consecutive results merge into one Anthropic user message
An assistant message with content: null and no tool_callsDropped and recorded (OpenAI's own refusal shape); it has nothing Anthropic can carry
stream_optionsUsage is on by default; an explicit stream_options.include_usage: false is honored — you are still billed, but no usage chunk is emitted to you
nRejected when n > 1: 400 invalid_request, on /v1/chat/completions and /v1/messages. The normalizer emits one choice; billing a garbled multi-choice response would be dishonest. The images route takes n up to 4 — see /docs/api-images
temperatureForwarded when ≤ 1. OpenAI's 0..2 range does not map to Anthropic's 0..1: temperature > 1 is a typed 400, never a silent clamp
top_pForwarded only when temperature is absent (Anthropic documents them as mutually exclusive; temperature wins, the drop is recorded)
parallel_tool_calls: false→ tool_choice.disable_parallel_tool_use: true
tools[].function.strictMapped 1:1
response_format, reasoning_effort, prediction, audio, modalitiesTyped 400 naming the param — the gateway does not map these across dialects yet and will not drop them silently: pinning JSON mode or a reasoning budget must not silently produce a different answer
Trailing assistant messageTyped 400: Anthropic treats it as an assistant-prefill request, which current Claude models reject
Everything else (logprobs, seed, penalties, …)Dropped and recorded — never forwarded

Anthropic surface → OpenAI-dialect provider

ParameterHandling
model, stream, temperature, top_pCopied verbatim
max_tokens→ max_completion_tokens (the modern OpenAI field; legacy max_tokens is rejected by o-series/GPT-5 reasoning models). Set to 4096 by the gateway when absent, as above
system (string or text blocks)→ one leading system message
stop_sequences→ stop
messagesTranslated: tool_use ↔ tool_calls, tool_result blocks → role:"tool" messages (one per result)
tools, tool_choiceTranslated: auto→"auto", any→"required", none→"none", {type:"tool",name}→{function:{name}}. tool_choice.disable_parallel_tool_use: true → parallel_tool_calls: false; tools[].strict maps 1:1. Server tools (computer use, web search, bash, text editor) are a typed 400 — an OpenAI-dialect deployment cannot execute them
thinking, output_configTyped 400 naming the param (not mappable across dialects yet; native passthrough on Anthropic-dialect routes)
top_k, metadataDropped and recorded (native passthrough on Anthropic-dialect routes). On a house route, metadata.user_id is used as the end-user id and is not recorded (see Sessions and prompt caching)
cache_control (nested and top-level)To OpenRouter: kept on the OpenAI text parts, and the top-level field too. A mark on a tool definition or a tool_use block has no place in OpenAI format: dropped and recorded, e.g. tools[0].cache_control. To OpenAI direct and Azure: dropped and recorded with the full path, e.g. messages[0].content[0].cache_control
tool_result.is_error (nested)Dropped and recorded with its full path
Note

"Dropped" means exactly that: the parameter is never forwarded — and every drop is recorded, per request. The full paths appear in the x-tm-dropped-params response header, in each attempt's dropped_params in GET /v1/generation, and in the (content-free) ledger. Send "provider": {"require_parameters": true} to turn any would-be drop into a typed 400 naming the first dropped parameter instead of a dispatch. Dropping only ever applies to parameters, never content.

Content is never silently dropped

Parameters are droppable because they don't get billed; content is not. A provider that drops a part still bills the request, so two rules apply:

  • A part goes only to a model that takes it. Each listing declares what a message to the model may carry: text, plus any of image, file (a document such as a PDF), audio and video. GET /v1/models shows what a model takes for your key in architecture.input_modalities. A model's page lists only the inputs that every listing of the model takes, so it can show less than the gateway accepts. The gateway sends a part only to a deployment whose listing takes its kind. A part that no deployment of the model takes is a typed 400 naming the exact field, on both surfaces, before anything is reserved or sent: messages[0].content[1]: model "deepseek/deepseek-v4-flash" does not take image input; it takes text. The parts checked are image_url, file, input_audio and video_url (or input_video) on the OpenAI surface, and image and document blocks on the Anthropic surface, also inside a tool_result. A route on your own provider key (BYOK) declares nothing, so the gateway sends every part to it as you sent it.
  • A translation names what it cannot carry. On a translated route, content the translation cannot represent — image parts, audio parts, unknown block types — is a typed 400 naming the exact field (messages[2].content[0] and so on), never a silent omission.

One deliberate exception: assistant thinking / redacted_thinking blocks echoed back on the Anthropic surface are dropped and recorded rather than rejected. The gateway's own surface emits those blocks (translated from upstream reasoning deltas), so echoing a transcript it produced must not 400 — but thinking content has no OpenAI-dialect equivalent and is never forwarded.

Sessions and prompt caching

A provider keeps a prompt's cache on one machine or host. The gateway passes on what each provider needs to send a conversation back there (decision D23).

Session id. The gateway takes the first valid value of these, in this order:

OrderInputWhere
1session_idBody
2x-session-idHeader
3x-session-affinityHeader
4x-claude-code-session-idHeader

A valid value has 1 to 256 printable ASCII characters. The body field prompt_cache_key follows the same rule. An invalid value is ignored, never an error: the gateway uses the next input and records the ignored one as session_id (invalid), prompt_cache_key (invalid) or header:<name> (invalid). With provider.require_parameters: true, an ignored input is a 400, as for any drop. Valid session inputs are never recorded as dropped.

Each route gets its own field:

RouteField sentValue
OpenAI (house or BYOK), Azureprompt_cache_keyYour prompt_cache_key, else the session id. Nothing if you sent neither
OpenRoutersession_idThe session id, else your prompt_cache_key, else an id that the gateway makes from your organization, the model and the opening messages. Your prompt_cache_key goes too
Anthropic, BedrockNothingAnthropic has no session input: its cache matches the prompt prefix

Values go unchanged. A value longer than the provider accepts is replaced by tm- followed by 43 base64url characters: a hash of the value. prompt_cache_retention ("in_memory" or "24h") goes unchanged to OpenAI and Azure; other routes drop and record it.

End-user ids on house routes. A house route runs on the marketplace's own provider account, which every customer shares. On a house route, the gateway sends exactly one end-user id: a hash of your organization and your safety_identifier (else user, else, on /v1/messages, metadata.user_id), or of your organization alone.

House routeField sentNot sent
openaisafety_identifieruser
openrouterusersafety_identifier
anthropicmetadata.user_iduser, safety_identifier

BYOK routes get your values unchanged: user and metadata as before, and safety_identifier on OpenAI and Azure from /v1/chat/completions. Anthropic and Bedrock BYOK routes, and any translation from /v1/messages, drop and record it.

Cache marks. The gateway never adds a cache_control mark that you did not send. The ledger prices every cache write at the 5-minute rate, so on a house route ttl: "1h" is removed from every mark (the cache then keeps 5 minutes) and recorded as <path>.cache_control.ttl, for example system[0].cache_control.ttl. ttl: "5m" passes. BYOK routes keep ttl: "1h": you pay that provider directly.

Staying on the first route. When the gateway's own quota pool for your first route is full, a follow-up request (one with an assistant or tool message after a user message) waits up to 5 seconds for it, instead of moving to a route with a cold cache. The response then carries x-tm-affinity-wait-ms with the milliseconds waited.

Usage accounting (how tokens are counted)

ConceptOpenAI surfaceAnthropic surface
Billable inputprompt_tokens — cache-inclusive: input + cache reads + cache writesinput_tokens — excludes cache reads and writes
Cache readsprompt_tokens_details.cached_tokenscache_read_input_tokens
Cache writesprompt_tokens_details.cache_write_tokenscache_creation_input_tokens
Outputcompletion_tokensoutput_tokens
Reasoningcompletion_tokens_details.reasoning_tokens (a subset of output)not separately reported

Cross-dialect translation converts exactly by this table; total_tokens = prompt_tokens + completion_tokens.

  • Usage is always on. Streams and non-streams both report it. The gateway injects stream_options.include_usage: true into every streamed OpenAI-dialect dispatch itself — without it those streams carry no usage at all. The only stream_options value that changes what you receive is an explicit include_usage: false, which suppresses the usage chunk (you are billed the same).
  • cache_write_tokens appears on both surfaces. Cache writes bill at their own rate, so the split is always visible: on billed OpenAI-surface responses prompt_tokens_details.cache_write_tokens is always present (0 when the provider reports none), stream and non-stream alike.
  • Every billed response carries usage.cost (USD) — in the final usage chunk on streams, in the body otherwise — computed with the exact integer micro-USD math the ledger settles with, and covering the full request debit, including attempts that failed over before your answer started. You can always recompute your bill from the wire.
  • Rounding favors you. Prices are micro-USD per million tokens; the pre-flight reservation rounds up, settlement rounds down.
Warning

One documented limitation: an Anthropic-surface stream that dies mid-answer has no legal wire slot for usage (message_delta is the only one, and it never arrives). Recompute those from GET https://api.routerplus.com/v1/generation?id=REQUEST_ID, which returns every physical attempt with tokens and settled cost. The OpenAI surface has no such gap — known usage and cost are emitted before the terminal error event.

Streams

OpenAI-surface event order: role-priming delta (sent as soon as the upstream proves alive) → content deltas → finish chunk → usage chunk → data: [DONE]. [DONE] is never sent after an error, and nothing ever follows a terminal event.

  • Keep-alives. After output has started, any silence of 15 seconds or more gets a keep-alive frame so proxies and clients don't kill an idle-but-healthy stream (slow reasoning models): an SSE comment (: processing) on the OpenAI surface, a native ping event on the Anthropic surface. Both are spec-ignorable — SDKs skip them without code changes.
  • The commit boundary. Before any semantic output reaches you, provider failures are handled by invisible zero-backoff failover — you see one response, one monotonic event sequence, and x-tm-attempts counts what it took. After output has started, the gateway never switches providers: an upstream death becomes one terminal error event inside the 200 stream, and retry semantics stay yours. A refusal is never rerouted to another provider.
  • The request deadline. 450 seconds from dispatch, or a routing policy's timeout_ms. A stream still open then gets the same terminal error event.
  • Mid-stream terminal shape: OpenAI surface — a full chunk envelope with finish_reason: "error" and a top-level error object; Anthropic surface — an event: error frame.

finish_reason and native_finish_reason

Translated responses map finish reasons both ways:

Anthropic stop_reasonOpenAI finish_reason
end_turnstop
stop_sequencestop
max_tokenslength
tool_usetool_calls
refusalcontent_filter

The mapping is lossy (two stop reasons both become stop), so translated responses on the OpenAI surface carry the provider's raw value verbatim in choices[0].native_finish_reason — on the non-stream body and on the stream's finish chunk. When it differs from finish_reason, the native value is what the provider actually said — trust it for analytics. Passthrough responses are relayed verbatim, so their finish_reason is already native.

Errors

Every error body is generated by the gateway in the dialect you called with, and carries a stable error_type with the canonical class — error.metadata.error_type on the OpenAI surface, error.error_type on the Anthropic surface — plus the dialect's native fields so your SDK's built-in handling works untouched. On the Anthropic surface that means real native type strings (billing_error, rate_limit_error, authentication_error, not_found_error, api_error, invalid_request_error), so Anthropic SDK retry classification behaves. A provider's own error body is never relayed; the provider's HTTP status rides in x-tm-upstream-status. The body's metadata also names where the failure came from (origin, and the limit that refused it when one did) — see Errors.

Canonical classes and how routing treats them:

Upstream signalerror_typeFailover
429 (Retry-After honored, capped at 60s)rate_limityes
401/402/403 from a provider (our account problem, not yours)authyes
Context/length errorscontext_overflowno — fix the request or pick a bigger window
Content policy / refusalcontent_policynever — a refusal is not rerouted
404 / model not found upstreammodel_unavailableyes
5xx / 529 / overloadedupstream_erroryes
Other 4xxupstream_errorno

On POST /v1/decisions some of these signals read differently, by their status and the provider's own error code, never by the provider's name: Levanto's 402, Fastino's 402 or 403, a System One provider's 401, 402 or 403 on our account, and Cloudflare's 429 with code 3036 (our free daily allocation is used up) are a 503 model_unavailable with metadata.retryable: false. Cloudflare's 413 with code 5021 is a 413 context_overflow; Perplexity's 400 that names the model's maximum context length (on Perplexity Decider v1 27B) is a 400 context_overflow, by the context rule above; and a System One provider's 408 is a retryable 504 upstream_error. A decisions provider's 429 never counts against the deployment's health circuit. See POST /v1/decisions.

Out-of-credits is deliberately a 429 insufficient_quota (not a 402): byte-compatible with what OpenAI's SDKs already recognize and back off on. Gateway-authored error bodies never echo flagged input content. The full per-status remediation table is on Errors.

Response headers

HeaderWhenMeaning
x-request-idevery callthe id GET /v1/generation?id= audits
x-tm-providerserved requestsdeployment that served the answer
x-tm-attemptsserved / failed dispatchesphysical dispatches, failovers included
x-tm-upstream-statuswhen an upstream respondedthe provider's own HTTP status
x-tm-error-codefailurescanonical class
x-tm-error-originlimit, provider and infrastructure failureswhere it came from: gateway_admission, upstream_quota, upstream, gateway_infrastructure or authorization
x-tm-limit-scope, x-tm-limit-kind, x-tm-limit-ida refused limitwhich limit refused the request — see Limits and capacity
x-tm-cap-resetspend-cap 429sexact instant the monthly cap resets (first of next UTC month)
x-tm-dropped-paramswhen non-emptycomma-joined names of parameters the gateway stripped (D8 §2: drops are recorded, never silent)
x-tm-upstream-modelaggregator swapsthe id the gateway actually sent upstream
x-tm-served-byaggregator routes that name itthe provider the aggregator used, e.g. Amazon Bedrock
x-tm-route-plan-idbilled requeststhe id of the route plan that chose the deployments
x-tm-admission-modebilled requestsshared, or bounded_local while the shared limit store is unreachable
x-tm-remaining-rpm, x-tm-remaining-tpmadmitted requeststhe smallest remaining allowance across the limits the request claimed
retry-after429/503seconds to wait; a provider's value is honored, capped at 60s

Browsers can read every header above except x-tm-error-code and x-tm-route-plan-id, which are not CORS-exposed; use the body's error_type in browser code.

Optional, content-free app attribution: send HTTP-Referer and X-Title to identify your app in per-app analytics.

Out of scope in v1

Honest errors or recorded drops — nothing silent:

  • Image, file, audio and video input have no cross-dialect translation: typed 400 on translated routes. On a passthrough route a part reaches the provider as sent when the model takes its kind, and is a typed 400 when no deployment of the model takes it; GET /v1/models lists the inputs a model takes.
  • Image and video generation are in scope on their own surfaces (D8 §7.6, D17, D18); image edits, variations and streaming, and image-to-video, are not — see POST /v1/images/generations and POST /v1/videos.
  • n > 1 choices: rejected 400 on the chat surfaces (the images route takes n up to 4).
  • No model id variants — ids containing : are not in the catalog and 404 like any unlisted id, and there is never a silent substitute.
  • On passthrough routes, the forwarded parameters reach the provider as you sent them, and the provider's behavior is the provider's. The guaranteed surface is exactly what this page documents.

See also: Migrate · Quickstart · Errors

Markdown source for agents: /docs/compat.md · index at /llms.txt