Errors
Every error class, per-surface shapes, exact remediation.
Every failure carries one canonical error class — a short stable string like rate_limit or insufficient_quota — in a fixed spot in the body and in the x-tm-error-code response header. The body's native shape follows the surface you called (OpenAI-compatible or Anthropic-compatible), so your SDK's built-in error handling keeps working; the canonical class is there so your code never has to parse prose.
A failure that happened after your request was accepted — at a limit, at a provider, or in our infrastructure — also says where it came from. The x-tm-error-origin header (and origin in the body's metadata) is one of gateway_admission (a limit or balance of yours), upstream_quota (a provider's declared capacity), upstream (the provider answered), gateway_infrastructure (our side) or authorization (your key or connection). When a limit refused the request, x-tm-limit-scope, x-tm-limit-kind and, where known, x-tm-limit-id name it. See Limits and capacity.
Where the class lives
OpenAI surface (POST /v1/chat/completions):
{
"error": {
"code": "model_unavailable",
"message": "model \"gpt-5-nano\" is not in the catalog; GET /v1/models lists what this key can serve",
"type": "invalid_request_error",
"metadata": { "error_type": "model_unavailable" }
},
"request_id": "6f8f57b2-..."
}Anthropic surface (POST /v1/messages):
{
"type": "error",
"error": {
"type": "not_found_error",
"message": "model \"gpt-5-nano\" is not in the catalog; GET /v1/models lists what this key can serve",
"error_type": "model_unavailable"
},
"request_id": "6f8f57b2-..."
}Four rules govern these bodies:
error.metadata.error_type(OpenAI surface) orerror.error_type(Anthropic surface) is always the canonical class. Every error body is generated by the gateway; a provider's own error body is never relayed. The provider's HTTP status is inx-tm-upstream-statuswhen it answered.- On the Anthropic surface,
error.typeadditionally maps to the native Anthropic type string (rate_limit_error,billing_error, ...) so the Anthropic SDK's own error classification and retry behavior work unmodified. - On the OpenAI surface,
error.codecarries the canonical class. Switch onerror.code,metadata.error_type, or the HTTP status — not onerror.type, which is a generic string. The one exception is out-of-credits: there bothcodeandtypeareinsufficient_quota, byte-matching OpenAI's own wire so OpenAI SDKs recognize it natively. - Errors raised before routing (bad key, unknown route, a crash in our handler) pick the shape from the path: requests to
/v1/messagesget Anthropic-shaped bodies, everything else OpenAI-shaped.
The metadata carries a few more fields after a limit or a balance refused the request: origin, limit_scope, limit_kind, limit_id, retryable, and retry_at or reset_at when one is known. On the OpenAI surface they sit beside error_type in error.metadata; on the Anthropic surface in error.metadata.
The canonical classes
error_type | HTTP | OpenAI error.code | Anthropic error.type | Meaning |
|---|---|---|---|---|
auth | 401 (or a relayed 401/402/403) | auth | authentication_error | your key is missing, wrong, or disabled — or, rarely, the marketplace's own provider account failed on every candidate (see below) |
invalid_request | 400 | invalid_request | invalid_request_error | body is not valid JSON, n>1, or a request no deployment can serve as sent — temperature > 1, response_format or thinking across a dialect boundary, unsupported content (images, audio) across a dialect boundary, or a content part the model does not take (an image, file, audio or video part that no route of the model declares) — the message names the field; a request-level credential or destination (api_key, base_url, connection_id, …) or an unknown provider control; on /v1/images/generations and /v1/videos: a value outside the model's accepted values (the message names the field), n outside 1–4, stream/partial_images, response_format: url, an image input on the videos route, a model on the wrong route — and an image or video model on a chat route (the message names the right route); on /v1/decisions: a question in another decision model's format (schema on Jev, Mercury Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider v1 27B or Sage, questions on GLiNER-2.5-Decide), a question kind that does not fit the state, a field the model does not take, two GLiNER results with one name, or, on Perplexity Decider v1 27B, an object with "type": "image_url" in the state or a question, because decision models take text and JSON (the message names the field) |
model_not_priced | 400 | model_not_priced | invalid_request_error | the model is listed but has no price row; we refuse rather than guess $0 |
context_overflow | provider's status (usually 400; 413 from Cloudflare on Clef and Clef-flash) | context_overflow | invalid_request_error | prompt exceeds the model's context window; on /v1/decisions, Cloudflare refuses a Clef or Clef-flash request that it estimates above 65,536 tokens (its 413, code 5021), and Perplexity refuses a Perplexity Decider v1 27B question whose prompt (the state and that question) is above 262,144 tokens (its 400); the message carries the provider's words |
content_policy | provider's status (usually 400) | content_policy | invalid_request_error | the provider refused the content; an OpenAI image moderation refusal (moderation_blocked) is relayed at the provider's own status (usually 400): never billed, never rerouted, never a health strike |
model_unavailable | 404 / 502 / 503 | model_unavailable | not_found_error | 404: not in the catalog (message echoes your requested id), or not served by your key's connection · 502: connection to the provider failed on every candidate · 503, on /v1/decisions: the provider refuses our account for capacity: Levanto's monthly allowance for Sage is used up (the message says Sage is out of capacity until the provider's next billing period), or Fastino refuses GLiNER-2.5-Decide for credit or billing, or Bespoke Labs refuses Bespoke Nimble v3 for credit (its 402), or Cloudflare refuses Clef or Clef-flash for our token or plan (its 401 or 403) or because our account's free daily allocation is used up (its 429 with code 3036); a 401, 402 or 403 from any System One provider on our key reads the same (TypeSafe or OpenRouter, on Jev or Mercury Decide, and RouterPlus, on Decider 2B and Kev 4B, too, and Perplexity, on Perplexity Decider v1 27B), and on a connection with your own key it keeps its status — the message says the model is out of capacity at the provider |
upstream_error (timeout) | 504 | upstream_error | api_error | the provider exceeded the gateway's wait: 20 s to first response headers on streams, the generation budget (450 s) on non-stream calls, 60 s to accept a video job, or 60 s for an answer on /v1/decisions (a GLiNER-2.5-Decide cold start can outlast it; a RouterPlus cold start, 15 to 20 s, fits inside it). On /v1/decisions a System One provider's own 408 (Cloudflare's timeout on Clef and Clef-flash) is this class too, and so is Perplexity's own 504 on Perplexity Decider v1 27B (it counts against the circuit, as every 5xx does). Retryable; a lone timeout never trips the health circuit |
not_found | 404 | not_found | not_found_error | no such route — or, on /v1/videos/{id}, no such video for your organization, or on /v1/generation, an id that is not a request id |
request_too_large | 413 | request_too_large | request_too_large | body over 10 MB (1 MB on POST /v1/videos and POST /v1/decisions) |
rate_limit | 429 | rate_limit | rate_limit_error | a rate limit refused the request: the key's RPM, one of the organization's shared limits, a provider pool's declared capacity or its cooldown after a provider 429, or the provider rate-limited every candidate. x-tm-limit-scope and x-tm-limit-kind say which, and x-tm-limit-id whether a pool limit was the pool's (pool:…) or your organization's share of it (pool-share:…). On /v1/decisions, Mercury Decide's pool is small (20 requests a minute in all, 15 per organization; 5 in flight, 3 per organization) and OpenRouter caps free requests per day across our account: both come back as this class, and a provider's 429 never opens the deployment's health circuit; Bespoke Nimble v3's pool takes 7 requests at once, 5 per organization, because Bespoke allows 8 at once for our whole account and one is kept for our test environment, and Bespoke's own 429 comes back as this class too; Clef and Clef-flash share one pool (200 requests a minute, 150 per organization; at most 8 at a time, 6 per organization), and Cloudflare's 429 when Workers AI is busy (code 3040) comes back as this class too, while its code 3036 is the 503 model_unavailable above; Decider 2B and Kev 4B share one RouterPlus pool (11,400 requests a minute, 59 in flight, 75 % per organization), so your organization's own limits bind first, and RouterPlus's own 429 (about 200 requests a second for the whole endpoint) comes back as this class too and pauses the pool for all three; Perplexity Decider v1 27B's pool takes 540 requests a minute (405 per organization) and 8 at once (6 per organization), its hold counts the state once per question, so a long state with many questions can pass a token limit on its own (x-tm-limit-kind: tpm: ask fewer questions per call), and Perplexity's own 429 (10 requests a second for our organization) comes back as this class too |
email_not_verified | 403 | email_not_verified | permission_error | the key's account signed up but has not opened its email verify link yet; every request is refused until it does, even one that would cost nothing |
insufficient_quota | 429 | insufficient_quota | billing_error | balance too low, or a monthly spend cap hit |
usage_limited | 429 | usage_limited | rate_limit_error | unusual usage was detected on an organization that has only trial credit, and this request was limited; retry-after says when to try again |
usage_restricted | 403 | usage_restricted | permission_error | unusual usage was detected on an organization that has only trial credit, and this kind of request is refused until support reviews the account |
upstream_error | 5xx, or provider's 4xx | upstream_error | api_error | provider failed and no alternate could serve; at 4xx the request itself was rejected; on the images route a 200 that is unparseable or carries no image is a 502 that bills nothing |
gateway_error | 400 / 500 / 502 / 503 / 504 | gateway_error | api_error | 400: an output bound the gateway refused — max_tokens outside 1–32,768, or max_tokens and max_completion_tokens both present and different (x-tm-limit-kind: output_bound) · 500: our bug · 502: the provider's image response exceeded the 32 MiB cap, or a 2xx our translator cannot represent; nothing billed · 503: no capacity right now — every deployment serving the model is cooling down (x-tm-limit-kind: health, retry-after: 5), the platform or this worker is at capacity, the shared limit store or the ledger journal is unavailable; nothing was sent to a provider and nothing was charged · 504: the request's routing deadline (450 s, or the policy's timeout_ms) passed before a provider could be tried |
When the status comes from a provider, the provider's own HTTP status is also exposed in the x-tm-upstream-status header.
What to do, per class
| Class | Retry? | Remediation |
|---|---|---|
auth (401 from us) | no | send the key as Authorization: Bearer tm_vk_... or x-api-key: tm_vk_... — both work on every endpoint; complete browser signup and create a key in the console at https://app.routerplus.com/console/keys |
auth (relayed 401/402/403) | no | nothing — this is our provider account, not yours; see the note below |
invalid_request | never blind-retry | fix what the message names, then resend |
model_not_priced | no | report it with the x-request-id — a listed model without a price is our data bug |
context_overflow | never blind-retry | shorten the input or pick a larger-context model; deliberately not failed over |
content_policy | never blind-retry | change the prompt; never rerouted to another provider |
model_unavailable (404) | no | pick an id from GET /v1/models (or the public list at https://app.routerplus.com/api/models.json) |
model_unavailable (502) | yes | one retry after ~2 s; every candidate's connection failed |
model_unavailable (503, decisions) | not soon | the provider's capacity comes back when its allowance or credit does (for Sage, at Levanto's next billing period; for Clef and Clef-flash after Cloudflare's code 3036, at 00:00 UTC); use another decision model meanwhile, in its own question format. metadata.retryable is false. While the provider refuses, most requests to that model get gateway_error (503, "all deployments cooling down") instead, because each refusal counts against the provider's circuit: treat that the same way on Sage, GLiNER-2.5-Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B and Perplexity Decider v1 27B (Cloudflare's code 3036 is a 429, and a 429 never counts) |
not_found | no | endpoints: /v1/chat/completions, /v1/messages, /v1/messages/count_tokens, /v1/images/generations (alias /images/generations), /v1/videos, /v1/videos/{id}, /v1/videos/{id}/content, /v1/decisions, /v1/models, /v1/usage, /v1/generation, /v1/limits, /v1/route |
request_too_large | no | trim the body under 10 MB (1 MB for a video or decisions request) — the images route also caps the provider response at 32 MiB, but that is a 502 gateway_error, not a 413 |
rate_limit | yes | wait retry-after, retry once — details in Rate limits & spend caps and Limits and capacity |
email_not_verified | no | open the verify link in the signup email; the key works within seconds of the click |
usage_limited | after retry-after | honor retry-after; if it keeps happening, contact support |
usage_restricted | no | contact support to have the account reviewed |
insufficient_quota | no | retrying cannot help: add credits, or raise the cap at https://app.routerplus.com/console/billing; on cap breaches the exact reset time is in the message and x-tm-cap-reset |
upstream_error (5xx) | yes | one retry after ~2 s; a failover already happened if one was possible |
upstream_error (4xx) | never blind-retry | the request itself was rejected — the message names the provider and its status; x-tm-upstream-status carries it |
gateway_error (400) | never blind-retry | send one output limit, between 1 and 32,768 |
gateway_error (500) | no | report with the x-request-id |
gateway_error (502, images) | yes, once | nothing was billed; lower n or the quality |
gateway_error (503) | yes | wait retry-after and resend — the request never reached a provider, so a retry cannot double-bill |
gateway_error (504) | yes | resend; the deadline passed while candidates were being tried |
Failure headers
| Header | When | Meaning |
|---|---|---|
x-request-id | always | quote it in any report; also the id for GET /v1/generation?id= |
x-tm-error-code | every failure | the canonical class |
x-tm-error-origin | limit, provider and infrastructure failures | gateway_admission, upstream_quota, upstream, gateway_infrastructure or authorization; a request refused for its own shape (bad JSON, n>1, an unknown model) carries none |
x-tm-limit-scope, x-tm-limit-kind, x-tm-limit-id | a refused limit | scope (platform, org, workspace, principal, key, model, pool, endpoint), the kind of limit (rpm, rps, rps_queue, tpm, concurrency, spend, health, …), and the limit's id when one exists |
x-tm-attempts | after ≥1 dispatch | physical provider attempts, failovers included |
x-tm-upstream-status | when a provider answered | the provider's own HTTP status |
retry-after | 429 / 503 | seconds to wait — a rate limit says when its window frees; provider-sent values are honored but capped at 60; the all-cooling-down 503 sends 5; the ledger-journal 503 sends 2 |
x-tm-cap-reset | spend-cap 429 only | exact ISO instant the cap resets (first instant of the next UTC month) |
x-tm-error-code is not CORS-exposed; browser code reads the class from the body.
Strict by design
Five refusals people trip on are deliberate — the alternative in each case is billing you for something wrong.
n>1 is a typed 400, not a degraded answer. The normalizer emits exactly one choice; accepting n: 4 and silently billing a garbled single-choice response would be dishonest.
curl -s https://api.routerplus.com/v1/chat/completions \
-H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
-d '{"model":"claude-haiku-4-5","n":4,"messages":[{"role":"user","content":"hi"}]}'{
"error": {
"code": "invalid_request",
"message": "n>1 is not supported in v1; request a single choice",
"type": "invalid_request_error",
"metadata": { "error_type": "invalid_request" }
},
"request_id": "..."
}Content a model does not take is a typed 400, not an answer that ignored it. A provider that does not read images can drop the image, answer the rest, and still bill the request. So an image, file, audio or video part goes only to a route whose listing takes that kind, and a part that no route of the model takes is refused before anything is reserved or sent. GET /v1/models lists what each model takes in architecture.input_modalities.
curl -s https://api.routerplus.com/v1/chat/completions \
-H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
-d '{"model":"deepseek/deepseek-v4-flash","messages":[{"role":"user","content":[{"type":"text","text":"What is in this picture?"},{"type":"image_url","image_url":{"url":"https://example.com/cat.png"}}]}]}'{
"error": {
"code": "invalid_request",
"message": "messages[0].content[1]: model \"deepseek/deepseek-v4-flash\" does not take image input; it takes text",
"type": "invalid_request_error",
"metadata": { "error_type": "invalid_request" }
},
"request_id": "..."
}No silent model substitution. An unknown model id is a 404 whose message echoes exactly the string you sent — we never quietly swap in a "close enough" model. The 404 is identical whether the id never existed or simply is not served, so the error is not an existence oracle.
A content-policy refusal never reroutes. A refusal is not shopped around to a more permissive provider — that is a hard routing rule with no exceptions. If you get content_policy, the fix is the prompt.
Images are all-or-nothing. A failed, over-cap or malformed generation is a 502 and bills nothing; n above 4 on the images route is a typed 400 like n > 1 on chat. See POST /v1/images/generations.
Failover, and what the error you see means
Before any output has reached you, the gateway fails over between deployments invisibly, at zero backoff. Classes that are failover-eligible: rate_limit, auth (upstream), model_unavailable, and 5xx upstream_error. Classes that never fail over: content_policy (never, as above), context_overflow (its own typed class — reshape the request instead), and provider 4xx rejections (the request is the problem). So if a failover-eligible class reaches you, every healthy candidate was tried — x-tm-attempts says how many.
A relayed 402 (class auth) is never your fault. An upstream 401/402/403 means the marketplace's account with that provider has a credential or payment problem. Payability is treated as a health signal: the error is failover-eligible and counts against that provider's uptime, so an unpayable provider drains traffic automatically. You would only ever see it when every deployment serving the model has the same problem. On /v1/decisions most of these refusals come back as 503 model_unavailable instead, with metadata.retryable: false (see the table above).
Mid-stream failures (HTTP 200 already committed)
Once your answer has started streaming, the gateway never switches providers — a provider death mid-answer becomes exactly one terminal error event inside the 200 stream, and nothing follows it (no data: [DONE], no keep-alives, no more chunks).
OpenAI surface — a real chunk envelope with finish_reason: "error" and a top-level error, so SDKs raise and naive parsers still terminate cleanly:
{
"id": "chatcmpl-...",
"object": "chat.completion.chunk",
"created": 1757000000,
"model": "claude-haiku-4-5",
"choices": [{ "index": 0, "delta": { "content": "" }, "finish_reason": "error" }],
"error": {
"code": "upstream_error",
"message": "upstream failure",
"type": "gateway_error",
"metadata": { "error_type": "upstream_error" }
}
}Anthropic surface — a native error event, with the canonical class riding alongside the native type:
{ "type": "error", "error": { "type": "api_error", "message": "upstream failure", "error_type": "upstream_error" } }On the OpenAI surface, any known usage and cost are emitted before the terminal error, so the wire stays billable-recomputable. The Anthropic wire has no legal usage slot outside message_delta; for an Anthropic-surface stream that died mid-answer, recompute from GET /v1/generation?id=<x-request-id>.
Retrying after a mid-stream error is your call — the answer never restarts on a different provider by itself.
Reading the class in code
import os
from openai import OpenAI, APIStatusError
client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])
try:
out = client.chat.completions.create(
model="claude-haiku-4-5",
messages=[{"role": "user", "content": "hello"}],
)
except APIStatusError as e:
err = e.response.json().get("error") or {}
kind = (err.get("metadata") or {}).get("error_type") \
or e.response.headers.get("x-tm-error-code")
origin = e.response.headers.get("x-tm-error-origin")
request_id = e.response.headers.get("x-request-id")
# kind is one of the canonical classes in the table aboveRetry discipline for agents
rate_limit→ waitretry-after, retry once.email_not_verified→ do not loop; ask the human to open the verify link in their email.insufficient_quota→ do not loop; surface to a human (money). See Rate limits & spend caps.usage_limited→ do not loop; honorretry-afteronce, then tell the human to contact support.usage_restricted→ do not retry; tell the human to contact support.upstream_errorat 5xx → one retry after ~2 s, then surface with thex-request-id. At 4xx the request itself is the problem — never blind-retry.upstream_errororgateway_errorat 502 on the images route → one retry; nothing was billed.gateway_errorat 503 or 504 → waitretry-after(or ~2 s), retry; nothing was billed.model_unavailableat 404 → pick a real id fromGET /v1/models; at 502 one retry.content_policy,context_overflow,invalid_request,gateway_errorat 400 → never blind-retry; the request must change first.
A video that fails
A video job ends failed without an HTTP error: POST /v1/videos answered 200 when the job was accepted, and the failure arrives later, on GET /v1/videos/{id}, as status: "failed" with error: {code, message}. The code is one of video_generation_failed, content_policy_violation, video_cancelled, video_expired or video_timeout. A failed job costs nothing; content_policy_violation means change the prompt. See POST /v1/videos.
Markdown source for agents: /docs/errors.md · index at /llms.txt