# Wire compatibility

This page is the published form of decision D8 — the conventions your bill is
computed from. Changing anything here is treated internally as a
billing-contract change and requires a recorded decision.

## Two surfaces, one catalog

| Surface | Endpoint | Auth |
|---|---|---|
| OpenAI-compatible | `POST https://api.routerplus.com/v1/chat/completions` | `Authorization: Bearer` or `x-api-key` |
| Anthropic-compatible | `POST https://api.routerplus.com/v1/messages` | `Authorization: Bearer` or `x-api-key` |
| OpenAI Images | `POST https://api.routerplus.com/v1/images/generations` | same |
| OpenAI Videos | `POST https://api.routerplus.com/v1/videos` | same |

Every cataloged chat model is callable from both chat surfaces; image models
answer on the Images surface only, video models on the Videos surface only.
`GET https://api.routerplus.com/v1/models` lists what your key can serve — OpenAI list
shape by default, Anthropic shape when you send an `anthropic-version` header.

One distinction drives everything below:

- **Passthrough route** — your dialect matches the serving provider's (OpenAI
  format → OpenAI-dialect provider, Anthropic format → Anthropic-dialect
  provider). The parameters your dialect defines are forwarded as you sent
  them, and the gateway adds `stream_options.include_usage` on streamed
  OpenAI-dialect dispatches. Whatever the provider accepts for those
  parameters, you get — including native features like Anthropic prompt
  caching (`cache_control`), `thinking`, `top_k`, and the `anthropic-beta`
  values listed below. A top-level key outside your dialect's parameter set is
  dropped and recorded, never forwarded. Session ids and cache marks follow
  [Sessions and prompt caching](#sessions-and-prompt-caching).
- **Translated route** — dialects differ, and the gateway translates the
  request, the stream, and errors. Translation covers a pinned parameter set;
  everything else follows the rules below.

A "typed 400" in the translation tables is decided per deployment. When a
translation refuses your request but another deployment of the model speaks
your own dialect, the request goes there instead; you get the 400 only when no
deployment can serve it. Content a model does not take follows the same rule:
a deployment that does not take a part is skipped, and the 400 comes only when
no deployment takes it (see
[Content is never silently dropped](#content-is-never-silently-dropped)).

## Parameter handling on passthrough routes

The forwarded set per dialect. Anything else at the top level is a recorded
drop (see the note under the translation tables).

| Dialect | Forwarded as sent |
|---|---|
| OpenAI → OpenAI-dialect provider | `model`, `messages`, `stream`, `stream_options`, `temperature`, `top_p`, `max_tokens`, `max_completion_tokens`, `stop`, `n`, `tools`, `tool_choice`, `parallel_tool_calls`, `response_format`, `reasoning_effort`, `seed`, `user`, `logprobs`, `top_logprobs`, `frequency_penalty`, `presence_penalty`, `logit_bias`, `metadata`, `store`, `prediction` |
| Anthropic → Anthropic-dialect provider | `model`, `messages`, `system`, `max_tokens`, `stream`, `temperature`, `top_p`, `top_k`, `stop_sequences`, `tools`, `tool_choice`, `metadata`, `thinking`, `output_config`, `cache_control` |

Some OpenAI-format keys go only to the providers that define them:

- `prompt_cache_retention` and `safety_identifier` go to OpenAI and Azure.
- The top-level `cache_control` (automatic caching) goes to OpenRouter.
- A `cache_control` mark on a content part passes unchanged to every OpenAI-dialect
  provider. OpenRouter uses it; OpenAI and Azure accept it and ignore it.

On Bedrock, the top-level `cache_control` becomes a mark on the last block that can
carry one, because Bedrock's API does not take the top-level field. On every route, a
top-level `cache_control` that Anthropic would refuse is dropped and recorded as
`cache_control`: a field other than `{"type": "ephemeral"}` with an optional `ttl` of `"5m"`
or `"1h"`, a fifth mark, a 1-hour field after a 5-minute mark, or a field whose TTL differs
from the mark on the last block.

Two Anthropic parameters are refused on every route, because stripping them
would change who runs what: `mcp_servers` and `container` are a typed 400.

The `anthropic-beta` header is forwarded to Anthropic-dialect providers only
for values that do not change how a token is billed: `prompt-caching`,
`token-efficient-tools`, `fine-grained-tool-streaming`, `interleaved-thinking`,
`claude-code`, `oauth` and `computer-use` prefixes. Any other value is dropped
and recorded as `header:anthropic-beta:<value>`.

## Parameter handling on translated routes

### OpenAI surface → Anthropic-dialect provider

| Parameter | Handling |
|---|---|
| `model`, `messages`, `stream`, `temperature`, `top_p` | Translated / copied verbatim |
| `max_tokens`, `max_completion_tokens` | Translated. The gateway sets `max_tokens: 4096` on any request that carries neither, before dispatch; see [POST /v1/chat/completions](/docs/api-chat-completions) for the 1–32,768 bound |
| `stop` (string or array) | → `stop_sequences` |
| `system` / `developer` messages | → the Anthropic `system` string, joined in order with blank lines. If a part has a `cache_control` mark, `system` becomes text blocks with the same text, and each mark ends a block |
| `cache_control` on a text part | Kept on the Anthropic block. A `role:"tool"` message's mark goes on its `tool_result` |
| Other keys of a text part | Not forwarded, and not recorded |
| `cache_control` (top-level) | → the Anthropic top-level field (automatic caching). On Bedrock it becomes a mark on the last block |
| `tools`, `tool_choice` | Translated: `auto`→`{type:"auto"}`, `required`→`{type:"any"}`, `{function:{name}}`→`{type:"tool",name}`, `none`→`{type:"none"}`. With `none` on a tool-free history, tools are omitted entirely so you don't pay for definitions you forbade using. A tool without `parameters` gets an empty object schema |
| `role:"tool"` messages | → `tool_result` blocks; consecutive results merge into one Anthropic user message |
| An assistant message with `content: null` and no `tool_calls` | Dropped and recorded (OpenAI's own refusal shape); it has nothing Anthropic can carry |
| `stream_options` | Usage is on by default; an explicit `stream_options.include_usage: false` is honored — you are still billed, but no usage chunk is emitted to you |
| `n` | **Rejected** when `n > 1`: 400 `invalid_request`, on `/v1/chat/completions` and `/v1/messages`. The normalizer emits one choice; billing a garbled multi-choice response would be dishonest. The images route takes `n` up to 4 — see [/docs/api-images](/docs/api-images) |
| `temperature` | Forwarded when ≤ 1. OpenAI's 0..2 range does not map to Anthropic's 0..1: `temperature > 1` is a typed 400, never a silent clamp |
| `top_p` | Forwarded only when `temperature` is absent (Anthropic documents them as mutually exclusive; `temperature` wins, the drop is recorded) |
| `parallel_tool_calls: false` | → `tool_choice.disable_parallel_tool_use: true` |
| `tools[].function.strict` | Mapped 1:1 |
| `response_format`, `reasoning_effort`, `prediction`, `audio`, `modalities` | **Typed 400** naming the param — the gateway does not map these across dialects yet and will not drop them silently: pinning JSON mode or a reasoning budget must not silently produce a different answer |
| Trailing `assistant` message | **Typed 400**: Anthropic treats it as an assistant-prefill request, which current Claude models reject |
| Everything else (`logprobs`, `seed`, penalties, …) | **Dropped and recorded** — never forwarded |

### Anthropic surface → OpenAI-dialect provider

| Parameter | Handling |
|---|---|
| `model`, `stream`, `temperature`, `top_p` | Copied verbatim |
| `max_tokens` | → `max_completion_tokens` (the modern OpenAI field; legacy `max_tokens` is rejected by o-series/GPT-5 reasoning models). Set to 4096 by the gateway when absent, as above |
| `system` (string or text blocks) | → one leading `system` message |
| `stop_sequences` | → `stop` |
| `messages` | Translated: `tool_use` ↔ `tool_calls`, `tool_result` blocks → `role:"tool"` messages (one per result) |
| `tools`, `tool_choice` | Translated: `auto`→`"auto"`, `any`→`"required"`, `none`→`"none"`, `{type:"tool",name}`→`{function:{name}}`. `tool_choice.disable_parallel_tool_use: true` → `parallel_tool_calls: false`; `tools[].strict` maps 1:1. Server tools (computer use, web search, bash, text editor) are a **typed 400** — an OpenAI-dialect deployment cannot execute them |
| `thinking`, `output_config` | **Typed 400** naming the param (not mappable across dialects yet; native passthrough on Anthropic-dialect routes) |
| `top_k`, `metadata` | **Dropped and recorded** (native passthrough on Anthropic-dialect routes). On a house route, `metadata.user_id` is used as the end-user id and is not recorded (see [Sessions and prompt caching](#sessions-and-prompt-caching)) |
| `cache_control` (nested and top-level) | To OpenRouter: kept on the OpenAI text parts, and the top-level field too. A mark on a tool definition or a `tool_use` block has no place in OpenAI format: **dropped and recorded**, e.g. `tools[0].cache_control`. To OpenAI direct and Azure: **dropped and recorded with the full path**, e.g. `messages[0].content[0].cache_control` |
| `tool_result.is_error` (nested) | **Dropped and recorded with its full path** |

> [!NOTE]
> "Dropped" means exactly that: the parameter is never forwarded — and every
> drop is **recorded**, per request. The full paths appear in the
> `x-tm-dropped-params` response header, in each attempt's `dropped_params` in
> [`GET /v1/generation`](/docs/api-usage), and in the (content-free) ledger.
> Send `"provider": {"require_parameters": true}` to turn any would-be drop
> into a typed 400 naming the first dropped parameter instead of a dispatch.
> Dropping only ever applies to *parameters*, never content.

### Content is never silently dropped

Parameters are droppable because they don't get billed; content is not. A
provider that drops a part still bills the request, so two rules apply:

- **A part goes only to a model that takes it.** Each listing declares what a
  message to the model may carry: text, plus any of image, file (a document such
  as a PDF), audio and video. `GET /v1/models` shows what a model takes for your
  key in `architecture.input_modalities`. A model's page lists only the inputs
  that every listing of the model takes, so it can show less than the gateway
  accepts. The gateway sends a part only to a deployment whose listing takes its
  kind. A part that no
  deployment of the model takes is a **typed 400 naming the exact field**, on
  both surfaces, before anything is reserved or sent:
  `messages[0].content[1]: model "deepseek/deepseek-v4-flash" does not take image input; it takes text`.
  The parts checked are `image_url`, `file`, `input_audio` and `video_url` (or
  `input_video`) on the OpenAI surface, and `image` and `document` blocks on the
  Anthropic surface, also inside a `tool_result`. A route on your own provider
  key (BYOK) declares nothing, so the gateway sends every part to it as you sent
  it.
- **A translation names what it cannot carry.** On a translated route, content
  the translation cannot represent — image parts, audio parts, unknown block
  types — is a **typed 400 naming the exact field** (`messages[2].content[0]`
  and so on), never a silent omission.

One deliberate exception: assistant `thinking` / `redacted_thinking` blocks
echoed back on the Anthropic surface are dropped and recorded rather than
rejected. The gateway's own surface emits those blocks (translated from
upstream reasoning deltas), so echoing a transcript it produced must not 400 —
but thinking content has no OpenAI-dialect equivalent and is never forwarded.

## Sessions and prompt caching

A provider keeps a prompt's cache on one machine or host. The gateway passes on what
each provider needs to send a conversation back there (decision D23).

**Session id.** The gateway takes the first valid value of these, in this order:

| Order | Input | Where |
|---|---|---|
| 1 | `session_id` | Body |
| 2 | `x-session-id` | Header |
| 3 | `x-session-affinity` | Header |
| 4 | `x-claude-code-session-id` | Header |

A valid value has 1 to 256 printable ASCII characters. The body field
`prompt_cache_key` follows the same rule. An invalid value is ignored, never an error:
the gateway uses the next input and records the ignored one as `session_id (invalid)`,
`prompt_cache_key (invalid)` or `header:<name> (invalid)`. With
`provider.require_parameters: true`, an ignored input is a 400, as for any drop. Valid
session inputs are never recorded as dropped.

Each route gets its own field:

| Route | Field sent | Value |
|---|---|---|
| OpenAI (house or BYOK), Azure | `prompt_cache_key` | Your `prompt_cache_key`, else the session id. Nothing if you sent neither |
| OpenRouter | `session_id` | The session id, else your `prompt_cache_key`, else an id that the gateway makes from your organization, the model and the opening messages. Your `prompt_cache_key` goes too |
| Anthropic, Bedrock | Nothing | Anthropic has no session input: its cache matches the prompt prefix |

Values go unchanged. A value longer than the provider accepts is replaced by `tm-`
followed by 43 base64url characters: a hash of the value. `prompt_cache_retention`
(`"in_memory"` or `"24h"`) goes unchanged to OpenAI and Azure; other routes drop and
record it.

**End-user ids on house routes.** A house route runs on the marketplace's own provider
account, which every customer shares. On a house route, the gateway sends exactly one
end-user id: a hash of your organization and your `safety_identifier` (else `user`,
else, on `/v1/messages`, `metadata.user_id`), or of your organization alone.

| House route | Field sent | Not sent |
|---|---|---|
| `openai` | `safety_identifier` | `user` |
| `openrouter` | `user` | `safety_identifier` |
| `anthropic` | `metadata.user_id` | `user`, `safety_identifier` |

BYOK routes get your values unchanged: `user` and `metadata` as before, and
`safety_identifier` on OpenAI and Azure from `/v1/chat/completions`. Anthropic and
Bedrock BYOK routes, and any translation from `/v1/messages`, drop and record it.

**Cache marks.** The gateway never adds a `cache_control` mark that you did not send.
The ledger prices every cache write at the 5-minute rate, so on a house route
`ttl: "1h"` is removed from every mark (the cache then keeps 5 minutes) and recorded as
`<path>.cache_control.ttl`, for example `system[0].cache_control.ttl`. `ttl: "5m"`
passes. BYOK routes keep `ttl: "1h"`: you pay that provider directly.

**Staying on the first route.** When the gateway's own quota pool for your first route
is full, a follow-up request (one with an assistant or tool message after a user
message) waits up to 5 seconds for it, instead of moving to a route with a cold cache.
The response then carries `x-tm-affinity-wait-ms` with the milliseconds waited.

## Usage accounting (how tokens are counted)

| Concept | OpenAI surface | Anthropic surface |
|---|---|---|
| Billable input | `prompt_tokens` — cache-inclusive: input + cache reads + cache writes | `input_tokens` — excludes cache reads and writes |
| Cache reads | `prompt_tokens_details.cached_tokens` | `cache_read_input_tokens` |
| Cache writes | `prompt_tokens_details.cache_write_tokens` | `cache_creation_input_tokens` |
| Output | `completion_tokens` | `output_tokens` |
| Reasoning | `completion_tokens_details.reasoning_tokens` (a subset of output) | not separately reported |

Cross-dialect translation converts exactly by this table;
`total_tokens = prompt_tokens + completion_tokens`.

- **Usage is always on.** Streams and non-streams both report it. The gateway
  injects `stream_options.include_usage: true` into every streamed
  OpenAI-dialect dispatch itself — without it those streams carry no usage at
  all. The only `stream_options` value that changes what you receive is an
  explicit `include_usage: false`, which suppresses the usage chunk (you are
  billed the same).
- **`cache_write_tokens` appears on both surfaces.** Cache writes bill at their
  own rate, so the split is always visible: on billed OpenAI-surface responses
  `prompt_tokens_details.cache_write_tokens` is always present (0 when the
  provider reports none), stream and non-stream alike.
- **Every billed response carries `usage.cost` (USD)** — in the final usage
  chunk on streams, in the body otherwise — computed with the exact integer
  micro-USD math the ledger settles with, and covering the *full* request
  debit, including attempts that failed over before your answer started. You
  can always recompute your bill from the wire.
- **Rounding favors you.** Prices are micro-USD per million tokens; the
  pre-flight reservation rounds up, settlement rounds down.

> [!WARNING]
> One documented limitation: an Anthropic-surface stream that dies mid-answer
> has no legal wire slot for usage (`message_delta` is the only one, and it
> never arrives). Recompute those from
> `GET https://api.routerplus.com/v1/generation?id=REQUEST_ID`, which returns every
> physical attempt with tokens and settled cost. The OpenAI surface has no such
> gap — known usage and cost are emitted *before* the terminal error event.

## Streams

OpenAI-surface event order: role-priming delta (sent as soon as the upstream
proves alive) → content deltas → finish chunk → usage chunk → `data: [DONE]`.
`[DONE]` is never sent after an error, and nothing ever follows a terminal
event.

- **Keep-alives.** After output has started, any silence of 15 seconds or more
  gets a keep-alive frame so proxies and clients don't kill an idle-but-healthy
  stream (slow reasoning models): an SSE comment (`: processing`) on the OpenAI
  surface, a native `ping` event on the Anthropic surface. Both are
  spec-ignorable — SDKs skip them without code changes.
- **The commit boundary.** Before any semantic output reaches you, provider
  failures are handled by invisible zero-backoff failover — you see one
  response, one monotonic event sequence, and `x-tm-attempts` counts what it
  took. After output has started, the gateway **never** switches providers: an
  upstream death becomes one terminal error event inside the 200 stream, and
  retry semantics stay yours. A refusal is never rerouted to another provider.
- **The request deadline.** 450 seconds from dispatch, or a routing policy's
  `timeout_ms`. A stream still open then gets the same terminal error event.
- Mid-stream terminal shape: OpenAI surface — a full chunk envelope with
  `finish_reason: "error"` and a top-level `error` object; Anthropic surface —
  an `event: error` frame.

## `finish_reason` and `native_finish_reason`

Translated responses map finish reasons both ways:

| Anthropic `stop_reason` | OpenAI `finish_reason` |
|---|---|
| `end_turn` | `stop` |
| `stop_sequence` | `stop` |
| `max_tokens` | `length` |
| `tool_use` | `tool_calls` |
| `refusal` | `content_filter` |

The mapping is lossy (two stop reasons both become `stop`), so translated
responses on the OpenAI surface carry the provider's raw value verbatim in
`choices[0].native_finish_reason` — on the non-stream body and on the stream's
finish chunk. When it differs from `finish_reason`, the native value is what
the provider actually said — trust it for analytics. Passthrough responses are
relayed verbatim, so their `finish_reason` is already native.

## Errors

Every error body is generated by the gateway in the dialect you called with,
and carries a stable `error_type` with the canonical class —
`error.metadata.error_type` on the OpenAI surface, `error.error_type` on the
Anthropic surface — plus the dialect's native fields so your SDK's built-in
handling works untouched. On the Anthropic surface that means real native type
strings (`billing_error`, `rate_limit_error`, `authentication_error`,
`not_found_error`, `api_error`, `invalid_request_error`), so Anthropic SDK
retry classification behaves. A provider's own error body is never relayed;
the provider's HTTP status rides in `x-tm-upstream-status`. The body's
`metadata` also names where the failure came from (`origin`, and the limit
that refused it when one did) — see [Errors](/docs/errors).

Canonical classes and how routing treats them:

| Upstream signal | `error_type` | Failover |
|---|---|---|
| 429 (Retry-After honored, capped at 60s) | `rate_limit` | yes |
| 401/402/403 from a provider (our account problem, not yours) | `auth` | yes |
| Context/length errors | `context_overflow` | no — fix the request or pick a bigger window |
| Content policy / refusal | `content_policy` | **never** — a refusal is not rerouted |
| 404 / model not found upstream | `model_unavailable` | yes |
| 5xx / 529 / overloaded | `upstream_error` | yes |
| Other 4xx | `upstream_error` | no |

On `POST /v1/decisions` some of these signals read differently, by their status and the
provider's own error code, never by the provider's name: Levanto's 402, Fastino's 402 or
403, a System One provider's 401, 402 or 403 on our account, and Cloudflare's 429 with
code 3036 (our free daily allocation is used up) are a 503 `model_unavailable` with
`metadata.retryable: false`. Cloudflare's 413 with code 5021 is a 413 `context_overflow`;
Perplexity's 400 that names the model's maximum context length (on Perplexity Decider v1
27B) is a 400 `context_overflow`, by the context rule above; and
a System One provider's 408 is a retryable 504 `upstream_error`. A decisions provider's 429
never counts against the deployment's health circuit. See
[POST /v1/decisions](/docs/api-decisions).

Out-of-credits is deliberately a 429 `insufficient_quota` (not a 402):
byte-compatible with what OpenAI's SDKs already recognize and back off on.
Gateway-authored error bodies never echo flagged input content. The full
per-status remediation table is on [Errors](/docs/errors).

## Response headers

| Header | When | Meaning |
|---|---|---|
| `x-request-id` | every call | the id `GET /v1/generation?id=` audits |
| `x-tm-provider` | served requests | deployment that served the answer |
| `x-tm-attempts` | served / failed dispatches | physical dispatches, failovers included |
| `x-tm-upstream-status` | when an upstream responded | the provider's own HTTP status |
| `x-tm-error-code` | failures | canonical class |
| `x-tm-error-origin` | limit, provider and infrastructure failures | where it came from: `gateway_admission`, `upstream_quota`, `upstream`, `gateway_infrastructure` or `authorization` |
| `x-tm-limit-scope`, `x-tm-limit-kind`, `x-tm-limit-id` | a refused limit | which limit refused the request — see [Limits and capacity](/docs/admission) |
| `x-tm-cap-reset` | spend-cap 429s | exact instant the monthly cap resets (first of next UTC month) |
| `x-tm-dropped-params` | when non-empty | comma-joined names of parameters the gateway stripped (D8 §2: drops are recorded, never silent) |
| `x-tm-upstream-model` | aggregator swaps | the id the gateway actually sent upstream |
| `x-tm-served-by` | aggregator routes that name it | the provider the aggregator used, e.g. `Amazon Bedrock` |
| `x-tm-route-plan-id` | billed requests | the id of the route plan that chose the deployments |
| `x-tm-admission-mode` | billed requests | `shared`, or `bounded_local` while the shared limit store is unreachable |
| `x-tm-remaining-rpm`, `x-tm-remaining-tpm` | admitted requests | the smallest remaining allowance across the limits the request claimed |
| `retry-after` | 429/503 | seconds to wait; a provider's value is honored, capped at 60s |

Browsers can read every header above except `x-tm-error-code` and
`x-tm-route-plan-id`, which are not CORS-exposed; use the body's `error_type`
in browser code.

Optional, content-free app attribution: send `HTTP-Referer` and `X-Title` to
identify your app in per-app analytics.

## Out of scope in v1

Honest errors or recorded drops — nothing silent:

- Image, file, audio and video **input** have no cross-dialect translation: typed
  400 on translated routes. On a passthrough route a part reaches the provider as
  sent when the model takes its kind, and is a typed 400 when no deployment of the
  model takes it; `GET /v1/models` lists the inputs a model takes.
- Image and video **generation** are in scope on their own surfaces (D8 §7.6,
  D17, D18); image edits, variations and streaming, and image-to-video, are
  not — see [POST /v1/images/generations](/docs/api-images) and
  [POST /v1/videos](/docs/api-videos).
- `n > 1` choices: rejected 400 on the chat surfaces (the images route takes `n` up to 4).
- No model id variants — ids containing `:` are not in the catalog and 404
  like any unlisted id, and there is never a silent substitute.
- On passthrough routes, the forwarded parameters reach the provider as you
  sent them, and the provider's behavior is the provider's. The guaranteed
  surface is exactly what this page documents.

See also: [Migrate](/docs/migration) · [Quickstart](/docs/quickstart) ·
[Errors](/docs/errors)
