Routing & failover
The commit boundary, no silent substitution, health circuits.
You ask for a model; the gateway picks a deployment that serves it and fails over between deployments when providers misbehave. Three guarantees shape everything on this page:
- 01The model you name is the model you get. An unknown model is an honest 404 naming the id — never a silent substitute.
- 02Failover only happens before your answer starts. Once the first token of real output reaches you, the request is committed to that provider forever.
- 03A content-policy refusal is never rerouted. Shopping a refused prompt around providers is a policy decision we refuse to make for you.
How a request picks a provider
The catalog is a list of deployments — each one a provider endpoint with a wire dialect, the model ids it serves, and a routing priority. Today most models have two: the lab's own API first, then OpenRouter as the fallback (anthropic then openrouter for Claude models, openai then openrouter for GPT models). Selection is a filter pipeline, entirely in-memory — routing itself makes no remote calls. (Billed requests separately consult the ledger and the shared admission store for authentication, limits and the reserve check.)
- 01Model filter — keep deployments whose model list contains the requested id, exactly. No fuzzy matching, no aliases. The one resolution: a dated snapshot of a listed id (
claude-haiku-4-5-20251001) routes and bills as its family, and your original id still goes to the provider verbatim. - 02Health filter — drop deployments whose circuit is open (cooling down).
- 03Priority sort — survivors are stable-sorted by ascending
priority; ties keep catalog order.
Dispatch then walks that ordered list. The first deployment to produce real output wins; failover-eligible failures move to the next candidate with zero backoff. Which dialect a deployment speaks is invisible to you — the gateway translates requests, streams, and errors both ways, so any listed model is callable from either surface.
A key with a BYOK connection or a routing policy runs the same pipeline over its own routes, with the policy's order, funding rule and attempt limit on top; POST /v1/route explains the plan for a request without dispatching it. See Routing policies.
A dedicated endpoint (<your-namespace>/<name>) routes to capacity reserved for your organization first, then to public deployments of the same model that meet the endpoint's constraints. It never routes to a different model. A closed endpoint names RouterPlus as the provider, whichever route served the request.
The commit boundary
Failover is governed by one line: the request commits to a deployment at its first semantic output — the first text delta, reasoning delta, tool call, refusal, or usage report that reaches you.
Before commit, failover is invisible. Failover is possible only while nothing has been written to you. A provider that fails before its stream produces a first frame — connection refused, a 5xx, a death before any bytes, or blowing the 20-second headers deadline — is skipped and the next candidate is tried immediately; your response headers are sent only once a live attempt starts writing. Once the headers and (on the OpenAI surface) the role-priming delta have gone out, the gateway no longer switches providers: a failure after that point becomes the terminal in-stream error described in Streaming, even if no semantic output was produced.
After commit, the gateway will never switch providers. Two different models do not produce interchangeable halves of one answer, and splicing them silently would be a lie about what you received. If the provider dies mid-answer you get the tokens it produced plus one terminal error event inside the 200 stream (exact frames in Streaming) — retrying is your decision, and the retry starts fresh.
You pay for every settled attempt, including ones discarded by pre-commit failover — a provider that dies after reporting prompt usage still billed us for that prompt, and pass-through pricing passes it through. The inline usage.cost on your response is the full request debit across all attempts, so the wire number always matches the ledger. x-tm-attempts tells you how many physical dispatches happened.
No silent model substitution
If the requested model is not in the catalog, the response is a 404 that names it:
{
"error": {
"code": "model_unavailable",
"message": "model \"gpt-5-ultra\" is not in the catalog; GET /v1/models lists what this key can serve",
"type": "invalid_request_error",
"metadata": { "error_type": "model_unavailable" }
},
"request_id": "…"
}There is no "closest match" fallback and no default model. The same honesty applies to requests we can't serve faithfully: n>1 is a typed 400 (the gateway normalizes every stream to a single choice, and billing you for a garbled multi-choice response would be worse than refusing).
What never fails over
Failover eligibility is decided once, at error-normalization time — routing never inspects provider-raw errors. Two classes are hard-excluded:
content_policy— checked before every status-code rule, so no HTTP status can override it: a refusal never retries and never reroutes to another provider. Error metadata on this path never echoes your flagged input. If you want a second opinion from a different provider, that is an explicit new request you make yourself.context_overflow— its own typed class. Another deployment of the same model has the same context window; the fix is reshaping the request or choosing a larger-context model, not blind rerouting.
Other buyer-fault 4xx errors (malformed request, etc.) also return directly: retrying an invalid request elsewhere just spends your money on the same error.
Failover eligibility by error class
| Upstream signal | error_type | Retryable | Fails over |
|---|---|---|---|
| 429 (Retry-After honored, capped 60s) | rate_limit | yes | yes |
| 401 / 402 / 403 from the provider | auth | no | yes |
| Context / length errors | context_overflow | no | no |
| Content policy / refusal shapes | content_policy | no | never |
| 404 / model_not_found upstream | model_unavailable | no | yes |
| 5xx / 529 / overloaded | upstream_error | yes | yes |
| Other 4xx | upstream_error | no | no |
A 401/402/403 from a provider is our account problem with that provider — not yours. It routes around the deployment and counts against that provider's health, so a provider we can no longer pay drains traffic automatically. You only see an auth error if every candidate failed.
Health circuits and cooldowns
Every deployment has an independent circuit breaker:
- Two consecutive counted failures open the circuit for a 30-second cooldown. While open, the deployment is skipped by routing.
- After the cooldown the circuit goes half-open: exactly one live request is admitted as a probe while everyone else keeps skipping. A successful probe closes the circuit; a failed probe re-opens it with a fresh cooldown. This avoids the stampede of a blind TTL expiry re-admitting all traffic at once.
- Only failover-eligible failures charge the circuit. Buyer-fault errors (400/413-class, content policy, context overflow) never do — your malformed request must not take a healthy provider out of rotation for everyone else. (An upstream 429 does charge the circuit — it is failover-eligible — but is excluded from uptime scoring as buyer-caused. It also cools that provider's admission pool for the provider's
retry-after, 1 to 60 seconds.) A client cancellation likewise counts in the provider's favor: you hung up; it did nothing wrong.
The same fairness rule shapes the public numbers: buyer-caused failures and cancellations are excluded from uptime scoring, and a model page shows no uptime percentage for a provider we call directly until it has 100 counted requests in the window — a 3-for-3 "100%" would be noise dressed up as a guarantee.
When everything is cooling down
If every deployment serving your model has an open circuit, the request is refused at once rather than queued: a 503 with retry-after: 5, x-tm-error-code: gateway_error, x-tm-limit-kind: health and the message all deployments cooling down. It is not a 404 — the model exists — and it is transient: a cooldown lasts 30 seconds, after which one request is let through as the probe. Retry after a few seconds.
Reading what happened
Every response tells you how it was served:
| Header | Meaning |
|---|---|
x-request-id | your handle for audit and support (always present) |
x-tm-provider | the deployment that served the request; routerplus on a closed dedicated endpoint |
x-tm-served-by | on an aggregator route, the provider that actually ran the model; absent on a closed dedicated endpoint |
x-tm-upstream-model | the id sent to the provider, when it differs from the one you sent; absent on a closed dedicated endpoint |
x-tm-attempts | physical dispatches, failovers included |
x-tm-upstream-status | the HTTP status the provider actually returned |
x-tm-dropped-params | request fields the gateway stripped before dispatch, named |
x-tm-route-plan-id | the plan id, for correlation with /v1/route |
x-tm-dedicated-endpoint, x-tm-route-role | on a dedicated endpoint: its id, and whether your dedicated capacity (primary) or the shared pool (fallback) served |
x-tm-queue-ms | on a dedicated endpoint with a paced limit: how long the request waited for its turn, in milliseconds |
x-tm-remaining-rpm, x-tm-remaining-tpm | what your rate limits had left at admission |
x-tm-error-code | canonical error class, on failures |
x-tm-error-origin | on failures: upstream, gateway_admission, upstream_quota, gateway_infrastructure or authorization |
retry-after | present on rate limits and cooldowns; capped at 60s |
On failures the body carries the stable canonical class — error.metadata.error_type on the OpenAI surface, error.error_type on the Anthropic surface — alongside each dialect's native type string so your SDK's built-in retry/backoff classification keeps working. The body is always the gateway's own shape; a provider's raw error body is never relayed. The full table is in Errors & remediation.
For the complete per-attempt story, the generation audit endpoint returns every physical attempt for a request id — which deployments were tried, in what order, and what each one cost:
curl -s -H "Authorization: Bearer $TM_API_KEY" \
"https://api.routerplus.com/v1/generation?id=<x-request-id>"A Claude request whose first attempt failed over from the lab's API to OpenRouter (trimmed: each attempt also carries its billing source, price snapshot, error origin, admission and route context, and dropped parameters):
{
"request_id": "2f6d8e3b-…",
"admission_events": [],
"attempts": [
{
"attempt": 1,
"deployment": "anthropic",
"model": "claude-sonnet-5",
"outcome": "failed",
"usage_provenance": "unknown",
"tokens": { "input": 0, "cached": 0, "cache_write": 0, "output": 0, "reasoning": 0 },
"reserved_max_usd": 0.04,
"cost_usd": 0,
"error_code": "upstream_error",
"dispatched_at": "2026-09-04T09:12:44.120Z"
},
{
"attempt": 2,
"deployment": "openrouter",
"model": "claude-sonnet-5",
"outcome": "completed",
"usage_provenance": "observed",
"tokens": { "input": 412, "cached": 0, "cache_write": 0, "output": 638, "reasoning": 0 },
"reserved_max_usd": 0.04,
"cost_usd": 0.007204,
"error_code": "",
"dispatched_at": "2026-09-04T09:12:45.010Z"
}
]
}Metadata only — prompts and responses are never stored. outcome is one of completed, failed, cancelled, incomplete (died after your answer started), or unknown_after_crash (settled at the reserved maximum until reconciled). admission_events lists refusals that happened before any attempt, such as a rate limit. This is also the recomputation path for the one wire gap described in Streaming: an Anthropic-surface stream that died mid-answer.
Markdown source for agents: /docs/routing.md · index at /llms.txt