# Routing & failover

You ask for a model; the gateway picks a deployment that serves it and fails
over between deployments when providers misbehave. Three guarantees shape
everything on this page:

1. **The model you name is the model you get.** An unknown model is an honest
   404 naming the id — never a silent substitute.
2. **Failover only happens before your answer starts.** Once the first token of
   real output reaches you, the request is committed to that provider forever.
3. **A content-policy refusal is never rerouted.** Shopping a refused prompt
   around providers is a policy decision we refuse to make for you.

## How a request picks a provider

The catalog is a list of **deployments** — each one a provider endpoint with a
wire dialect, the model ids it serves, and a routing `priority`. Today most
models have two: the lab's own API first, then OpenRouter as the fallback
(`anthropic` then `openrouter` for Claude models, `openai` then `openrouter` for
GPT models). Selection is a filter pipeline, entirely in-memory — routing itself
makes no remote calls. (Billed requests separately consult the ledger and the
shared admission store for authentication, limits and the reserve check.)

1. **Model filter** — keep deployments whose model list contains the requested
   id, exactly. No fuzzy matching, no aliases. The one resolution: a dated
   snapshot of a listed id (`claude-haiku-4-5-20251001`) routes and bills as its
   family, and your original id still goes to the provider verbatim.
2. **Health filter** — drop deployments whose circuit is open (cooling down).
3. **Priority sort** — survivors are stable-sorted by ascending `priority`;
   ties keep catalog order.

Dispatch then walks that ordered list. The first deployment to produce real
output wins; failover-eligible failures move to the next candidate with zero
backoff. Which dialect a deployment speaks is invisible to you — the gateway
translates requests, streams, and errors both ways, so any listed model is
callable from either surface.

A key with a BYOK connection or a routing policy runs the same pipeline over
its own routes, with the policy's order, funding rule and attempt limit on top;
`POST /v1/route` explains the plan for a request without dispatching it. See
[Routing policies](/docs/routing-policies).

A [dedicated endpoint](/docs/dedicated-endpoints) (`<your-namespace>/<name>`) routes to
capacity reserved for your organization first, then to public deployments of
the same model that meet the endpoint's constraints. It never routes to a
different model. A [closed](/docs/dedicated-endpoints#closed-endpoints) endpoint
names RouterPlus as the provider, whichever route served the request.

## The commit boundary

Failover is governed by one line: **the request commits to a deployment at its
first semantic output** — the first text delta, reasoning delta, tool call,
refusal, or usage report that reaches you.

**Before commit**, failover is invisible. Failover is possible only while
nothing has been written to you. A provider that fails before its stream
produces a first frame — connection refused, a 5xx, a death before any bytes,
or blowing the 20-second headers deadline — is skipped and the next candidate
is tried immediately; your response headers are sent only once a live attempt
starts writing. Once the headers and (on the OpenAI surface) the role-priming
delta have gone out, the gateway no longer switches providers: a failure after
that point becomes the terminal in-stream error described in
[Streaming](/docs/streaming), even if no semantic output was produced.

**After commit**, the gateway will never switch providers. Two different models
do not produce interchangeable halves of one answer, and splicing them silently
would be a lie about what you received. If the provider dies mid-answer you get
the tokens it produced plus one terminal error event inside the 200 stream
(exact frames in [Streaming](/docs/streaming)) — retrying is your decision, and
the retry starts fresh.

> [!NOTE]
> You pay for every settled attempt, including ones discarded by pre-commit
> failover — a provider that dies after reporting prompt usage still billed us
> for that prompt, and pass-through pricing passes it through. The inline
> `usage.cost` on your response is the **full request debit** across all
> attempts, so the wire number always matches the ledger. `x-tm-attempts` tells
> you how many physical dispatches happened.

## No silent model substitution

If the requested model is not in the catalog, the response is a 404 that names
it:

```json
{
  "error": {
    "code": "model_unavailable",
    "message": "model \"gpt-5-ultra\" is not in the catalog; GET /v1/models lists what this key can serve",
    "type": "invalid_request_error",
    "metadata": { "error_type": "model_unavailable" }
  },
  "request_id": "…"
}
```

There is no "closest match" fallback and no default model. The same honesty
applies to requests we can't serve faithfully: `n>1` is a typed 400 (the
gateway normalizes every stream to a single choice, and billing you for a
garbled multi-choice response would be worse than refusing).

## What never fails over

Failover eligibility is decided once, at error-normalization time — routing
never inspects provider-raw errors. Two classes are hard-excluded:

- **`content_policy`** — checked before every status-code rule, so no HTTP
  status can override it: a refusal never retries and **never** reroutes to
  another provider. Error metadata on this path never echoes your flagged
  input. If you want a second opinion from a different provider, that is an
  explicit new request you make yourself.
- **`context_overflow`** — its own typed class. Another deployment of the same
  model has the same context window; the fix is reshaping the request or
  choosing a larger-context model, not blind rerouting.

Other buyer-fault 4xx errors (malformed request, etc.) also return directly:
retrying an invalid request elsewhere just spends your money on the same error.

## Failover eligibility by error class

| Upstream signal | `error_type` | Retryable | Fails over |
|---|---|---|---|
| 429 (Retry-After honored, capped 60s) | `rate_limit` | yes | yes |
| 401 / 402 / 403 from the provider | `auth` | no | **yes** |
| Context / length errors | `context_overflow` | no | no |
| Content policy / refusal shapes | `content_policy` | no | **never** |
| 404 / model_not_found upstream | `model_unavailable` | no | yes |
| 5xx / 529 / overloaded | `upstream_error` | yes | yes |
| Other 4xx | `upstream_error` | no | no |

> [!NOTE]
> A 401/402/403 **from a provider** is our account problem with that provider —
> not yours. It routes around the deployment and counts against that provider's
> health, so a provider we can no longer pay drains traffic automatically. You
> only see an `auth` error if every candidate failed.

## Health circuits and cooldowns

Every deployment has an independent circuit breaker:

- **Two consecutive counted failures open the circuit** for a 30-second
  cooldown. While open, the deployment is skipped by routing.
- **After the cooldown the circuit goes half-open**: exactly one live request
  is admitted as a probe while everyone else keeps skipping. A successful probe
  closes the circuit; a failed probe re-opens it with a fresh cooldown. This
  avoids the stampede of a blind TTL expiry re-admitting all traffic at once.
- **Only failover-eligible failures charge the circuit.** Buyer-fault errors
  (400/413-class, content policy, context overflow) never do — your malformed
  request must not take a healthy provider out of rotation for everyone else.
  (An upstream 429 does charge the circuit — it is failover-eligible — but is
  excluded from uptime scoring as buyer-caused. It also cools that provider's
  admission pool for the provider's `retry-after`, 1 to 60 seconds.) A client
  cancellation likewise counts in the provider's favor: you hung up; it did
  nothing wrong.

The same fairness rule shapes the public numbers: buyer-caused failures and
cancellations are excluded from uptime scoring, and a model page shows no
uptime percentage for a provider we call directly until it has 100 counted
requests in the window — a 3-for-3 "100%" would be noise dressed up as a
guarantee.

## When everything is cooling down

If every deployment serving your model has an open circuit, the request is
refused at once rather than queued: a `503` with `retry-after: 5`,
`x-tm-error-code: gateway_error`, `x-tm-limit-kind: health` and the message
`all deployments cooling down`. It is not a 404 — the model exists — and it is
transient: a cooldown lasts 30 seconds, after which one request is let through
as the probe. Retry after a few seconds.

## Reading what happened

Every response tells you how it was served:

| Header | Meaning |
|---|---|
| `x-request-id` | your handle for audit and support (always present) |
| `x-tm-provider` | the deployment that served the request; `routerplus` on a closed dedicated endpoint |
| `x-tm-served-by` | on an aggregator route, the provider that actually ran the model; absent on a closed dedicated endpoint |
| `x-tm-upstream-model` | the id sent to the provider, when it differs from the one you sent; absent on a closed dedicated endpoint |
| `x-tm-attempts` | physical dispatches, failovers included |
| `x-tm-upstream-status` | the HTTP status the provider actually returned |
| `x-tm-dropped-params` | request fields the gateway stripped before dispatch, named |
| `x-tm-route-plan-id` | the plan id, for correlation with `/v1/route` |
| `x-tm-dedicated-endpoint`, `x-tm-route-role` | on a dedicated endpoint: its id, and whether your dedicated capacity (`primary`) or the shared pool (`fallback`) served |
| `x-tm-queue-ms` | on a dedicated endpoint with a [paced limit](/docs/dedicated-endpoints#the-paced-limit): how long the request waited for its turn, in milliseconds |
| `x-tm-remaining-rpm`, `x-tm-remaining-tpm` | what your rate limits had left at admission |
| `x-tm-error-code` | canonical error class, on failures |
| `x-tm-error-origin` | on failures: `upstream`, `gateway_admission`, `upstream_quota`, `gateway_infrastructure` or `authorization` |
| `retry-after` | present on rate limits and cooldowns; capped at 60s |

On failures the body carries the stable canonical class —
`error.metadata.error_type` on the OpenAI surface, `error.error_type` on the
Anthropic surface — alongside each dialect's native type string so your SDK's
built-in retry/backoff classification keeps working. The body is always the
gateway's own shape; a provider's raw error body is never relayed. The full
table is in [Errors & remediation](/docs/errors).

For the complete per-attempt story, the generation audit endpoint returns every
physical attempt for a request id — which deployments were tried, in what
order, and what each one cost:

```bash
curl -s -H "Authorization: Bearer $TM_API_KEY" \
  "https://api.routerplus.com/v1/generation?id=<x-request-id>"
```

A Claude request whose first attempt failed over from the lab's API to
OpenRouter (trimmed: each attempt also carries its billing source, price
snapshot, error origin, admission and route context, and dropped parameters):

```json
{
  "request_id": "2f6d8e3b-…",
  "admission_events": [],
  "attempts": [
    {
      "attempt": 1,
      "deployment": "anthropic",
      "model": "claude-sonnet-5",
      "outcome": "failed",
      "usage_provenance": "unknown",
      "tokens": { "input": 0, "cached": 0, "cache_write": 0, "output": 0, "reasoning": 0 },
      "reserved_max_usd": 0.04,
      "cost_usd": 0,
      "error_code": "upstream_error",
      "dispatched_at": "2026-09-04T09:12:44.120Z"
    },
    {
      "attempt": 2,
      "deployment": "openrouter",
      "model": "claude-sonnet-5",
      "outcome": "completed",
      "usage_provenance": "observed",
      "tokens": { "input": 412, "cached": 0, "cache_write": 0, "output": 638, "reasoning": 0 },
      "reserved_max_usd": 0.04,
      "cost_usd": 0.007204,
      "error_code": "",
      "dispatched_at": "2026-09-04T09:12:45.010Z"
    }
  ]
}
```

Metadata only — prompts and responses are never stored. `outcome` is one of
`completed`, `failed`, `cancelled`, `incomplete` (died after your answer
started), or `unknown_after_crash` (settled at the reserved maximum until
reconciled). `admission_events` lists refusals that happened before any
attempt, such as a rate limit. This is also the recomputation path for the one
wire gap described in [Streaming](/docs/streaming): an Anthropic-surface stream
that died mid-answer.
