Console
Core concepts/Errors

Errors

Every error class, per-surface shapes, exact remediation.

/llms.txt

Every failure carries one canonical error class — a short stable string like rate_limit or insufficient_quota — in a fixed spot in the body and in the x-tm-error-code response header. The body's native shape follows the surface you called (OpenAI-compatible or Anthropic-compatible), so your SDK's built-in error handling keeps working; the canonical class is there so your code never has to parse prose.

A failure that happened after your request was accepted — at a limit, at a provider, or in our infrastructure — also says where it came from. The x-tm-error-origin header (and origin in the body's metadata) is one of gateway_admission (a limit or balance of yours), upstream_quota (a provider's declared capacity), upstream (the provider answered), gateway_infrastructure (our side) or authorization (your key or connection). When a limit refused the request, x-tm-limit-scope, x-tm-limit-kind and, where known, x-tm-limit-id name it. See Limits and capacity.

Where the class lives

OpenAI surface (POST /v1/chat/completions):

json
{
  "error": {
    "code": "model_unavailable",
    "message": "model \"gpt-5-nano\" is not in the catalog; GET /v1/models lists what this key can serve",
    "type": "invalid_request_error",
    "metadata": { "error_type": "model_unavailable" }
  },
  "request_id": "6f8f57b2-..."
}

Anthropic surface (POST /v1/messages):

json
{
  "type": "error",
  "error": {
    "type": "not_found_error",
    "message": "model \"gpt-5-nano\" is not in the catalog; GET /v1/models lists what this key can serve",
    "error_type": "model_unavailable"
  },
  "request_id": "6f8f57b2-..."
}

Four rules govern these bodies:

  • error.metadata.error_type (OpenAI surface) or error.error_type (Anthropic surface) is always the canonical class. Every error body is generated by the gateway; a provider's own error body is never relayed. The provider's HTTP status is in x-tm-upstream-status when it answered.
  • On the Anthropic surface, error.type additionally maps to the native Anthropic type string (rate_limit_error, billing_error, ...) so the Anthropic SDK's own error classification and retry behavior work unmodified.
  • On the OpenAI surface, error.code carries the canonical class. Switch on error.code, metadata.error_type, or the HTTP status — not on error.type, which is a generic string. The one exception is out-of-credits: there both code and type are insufficient_quota, byte-matching OpenAI's own wire so OpenAI SDKs recognize it natively.
  • Errors raised before routing (bad key, unknown route, a crash in our handler) pick the shape from the path: requests to /v1/messages get Anthropic-shaped bodies, everything else OpenAI-shaped.

The metadata carries a few more fields after a limit or a balance refused the request: origin, limit_scope, limit_kind, limit_id, retryable, and retry_at or reset_at when one is known. On the OpenAI surface they sit beside error_type in error.metadata; on the Anthropic surface in error.metadata.

The canonical classes

error_typeHTTPOpenAI error.codeAnthropic error.typeMeaning
auth401 (or a relayed 401/402/403)authauthentication_erroryour key is missing, wrong, or disabled — or, rarely, the marketplace's own provider account failed on every candidate (see below)
invalid_request400invalid_requestinvalid_request_errorbody is not valid JSON, n>1, or a request no deployment can serve as sent — temperature > 1, response_format or thinking across a dialect boundary, unsupported content (images, audio) across a dialect boundary, or a content part the model does not take (an image, file, audio or video part that no route of the model declares) — the message names the field; a request-level credential or destination (api_key, base_url, connection_id, …) or an unknown provider control; on /v1/images/generations and /v1/videos: a value outside the model's accepted values (the message names the field), n outside 1–4, stream/partial_images, response_format: url, an image input on the videos route, a model on the wrong route — and an image or video model on a chat route (the message names the right route); on /v1/decisions: a question in another decision model's format (schema on Jev, Mercury Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider v1 27B or Sage, questions on GLiNER-2.5-Decide), a question kind that does not fit the state, a field the model does not take, two GLiNER results with one name, or, on Perplexity Decider v1 27B, an object with "type": "image_url" in the state or a question, because decision models take text and JSON (the message names the field)
model_not_priced400model_not_pricedinvalid_request_errorthe model is listed but has no price row; we refuse rather than guess $0
context_overflowprovider's status (usually 400; 413 from Cloudflare on Clef and Clef-flash)context_overflowinvalid_request_errorprompt exceeds the model's context window; on /v1/decisions, Cloudflare refuses a Clef or Clef-flash request that it estimates above 65,536 tokens (its 413, code 5021), and Perplexity refuses a Perplexity Decider v1 27B question whose prompt (the state and that question) is above 262,144 tokens (its 400); the message carries the provider's words
content_policyprovider's status (usually 400)content_policyinvalid_request_errorthe provider refused the content; an OpenAI image moderation refusal (moderation_blocked) is relayed at the provider's own status (usually 400): never billed, never rerouted, never a health strike
model_unavailable404 / 502 / 503model_unavailablenot_found_error404: not in the catalog (message echoes your requested id), or not served by your key's connection · 502: connection to the provider failed on every candidate · 503, on /v1/decisions: the provider refuses our account for capacity: Levanto's monthly allowance for Sage is used up (the message says Sage is out of capacity until the provider's next billing period), or Fastino refuses GLiNER-2.5-Decide for credit or billing, or Bespoke Labs refuses Bespoke Nimble v3 for credit (its 402), or Cloudflare refuses Clef or Clef-flash for our token or plan (its 401 or 403) or because our account's free daily allocation is used up (its 429 with code 3036); a 401, 402 or 403 from any System One provider on our key reads the same (TypeSafe or OpenRouter, on Jev or Mercury Decide, and RouterPlus, on Decider 2B and Kev 4B, too, and Perplexity, on Perplexity Decider v1 27B), and on a connection with your own key it keeps its status — the message says the model is out of capacity at the provider
upstream_error (timeout)504upstream_errorapi_errorthe provider exceeded the gateway's wait: 20 s to first response headers on streams, the generation budget (450 s) on non-stream calls, 60 s to accept a video job, or 60 s for an answer on /v1/decisions (a GLiNER-2.5-Decide cold start can outlast it; a RouterPlus cold start, 15 to 20 s, fits inside it). On /v1/decisions a System One provider's own 408 (Cloudflare's timeout on Clef and Clef-flash) is this class too, and so is Perplexity's own 504 on Perplexity Decider v1 27B (it counts against the circuit, as every 5xx does). Retryable; a lone timeout never trips the health circuit
not_found404not_foundnot_found_errorno such route — or, on /v1/videos/{id}, no such video for your organization, or on /v1/generation, an id that is not a request id
request_too_large413request_too_largerequest_too_largebody over 10 MB (1 MB on POST /v1/videos and POST /v1/decisions)
rate_limit429rate_limitrate_limit_errora rate limit refused the request: the key's RPM, one of the organization's shared limits, a provider pool's declared capacity or its cooldown after a provider 429, or the provider rate-limited every candidate. x-tm-limit-scope and x-tm-limit-kind say which, and x-tm-limit-id whether a pool limit was the pool's (pool:…) or your organization's share of it (pool-share:…). On /v1/decisions, Mercury Decide's pool is small (20 requests a minute in all, 15 per organization; 5 in flight, 3 per organization) and OpenRouter caps free requests per day across our account: both come back as this class, and a provider's 429 never opens the deployment's health circuit; Bespoke Nimble v3's pool takes 7 requests at once, 5 per organization, because Bespoke allows 8 at once for our whole account and one is kept for our test environment, and Bespoke's own 429 comes back as this class too; Clef and Clef-flash share one pool (200 requests a minute, 150 per organization; at most 8 at a time, 6 per organization), and Cloudflare's 429 when Workers AI is busy (code 3040) comes back as this class too, while its code 3036 is the 503 model_unavailable above; Decider 2B and Kev 4B share one RouterPlus pool (11,400 requests a minute, 59 in flight, 75 % per organization), so your organization's own limits bind first, and RouterPlus's own 429 (about 200 requests a second for the whole endpoint) comes back as this class too and pauses the pool for all three; Perplexity Decider v1 27B's pool takes 540 requests a minute (405 per organization) and 8 at once (6 per organization), its hold counts the state once per question, so a long state with many questions can pass a token limit on its own (x-tm-limit-kind: tpm: ask fewer questions per call), and Perplexity's own 429 (10 requests a second for our organization) comes back as this class too
email_not_verified403email_not_verifiedpermission_errorthe key's account signed up but has not opened its email verify link yet; every request is refused until it does, even one that would cost nothing
insufficient_quota429insufficient_quotabilling_errorbalance too low, or a monthly spend cap hit
usage_limited429usage_limitedrate_limit_errorunusual usage was detected on an organization that has only trial credit, and this request was limited; retry-after says when to try again
usage_restricted403usage_restrictedpermission_errorunusual usage was detected on an organization that has only trial credit, and this kind of request is refused until support reviews the account
upstream_error5xx, or provider's 4xxupstream_errorapi_errorprovider failed and no alternate could serve; at 4xx the request itself was rejected; on the images route a 200 that is unparseable or carries no image is a 502 that bills nothing
gateway_error400 / 500 / 502 / 503 / 504gateway_errorapi_error400: an output bound the gateway refused — max_tokens outside 1–32,768, or max_tokens and max_completion_tokens both present and different (x-tm-limit-kind: output_bound) · 500: our bug · 502: the provider's image response exceeded the 32 MiB cap, or a 2xx our translator cannot represent; nothing billed · 503: no capacity right now — every deployment serving the model is cooling down (x-tm-limit-kind: health, retry-after: 5), the platform or this worker is at capacity, the shared limit store or the ledger journal is unavailable; nothing was sent to a provider and nothing was charged · 504: the request's routing deadline (450 s, or the policy's timeout_ms) passed before a provider could be tried

When the status comes from a provider, the provider's own HTTP status is also exposed in the x-tm-upstream-status header.

What to do, per class

ClassRetry?Remediation
auth (401 from us)nosend the key as Authorization: Bearer tm_vk_... or x-api-key: tm_vk_... — both work on every endpoint; complete browser signup and create a key in the console at https://app.routerplus.com/console/keys
auth (relayed 401/402/403)nonothing — this is our provider account, not yours; see the note below
invalid_requestnever blind-retryfix what the message names, then resend
model_not_pricednoreport it with the x-request-id — a listed model without a price is our data bug
context_overflownever blind-retryshorten the input or pick a larger-context model; deliberately not failed over
content_policynever blind-retrychange the prompt; never rerouted to another provider
model_unavailable (404)nopick an id from GET /v1/models (or the public list at https://app.routerplus.com/api/models.json)
model_unavailable (502)yesone retry after ~2 s; every candidate's connection failed
model_unavailable (503, decisions)not soonthe provider's capacity comes back when its allowance or credit does (for Sage, at Levanto's next billing period; for Clef and Clef-flash after Cloudflare's code 3036, at 00:00 UTC); use another decision model meanwhile, in its own question format. metadata.retryable is false. While the provider refuses, most requests to that model get gateway_error (503, "all deployments cooling down") instead, because each refusal counts against the provider's circuit: treat that the same way on Sage, GLiNER-2.5-Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B and Perplexity Decider v1 27B (Cloudflare's code 3036 is a 429, and a 429 never counts)
not_foundnoendpoints: /v1/chat/completions, /v1/messages, /v1/messages/count_tokens, /v1/images/generations (alias /images/generations), /v1/videos, /v1/videos/{id}, /v1/videos/{id}/content, /v1/decisions, /v1/models, /v1/usage, /v1/generation, /v1/limits, /v1/route
request_too_largenotrim the body under 10 MB (1 MB for a video or decisions request) — the images route also caps the provider response at 32 MiB, but that is a 502 gateway_error, not a 413
rate_limityeswait retry-after, retry once — details in Rate limits & spend caps and Limits and capacity
email_not_verifiednoopen the verify link in the signup email; the key works within seconds of the click
usage_limitedafter retry-afterhonor retry-after; if it keeps happening, contact support
usage_restrictednocontact support to have the account reviewed
insufficient_quotanoretrying cannot help: add credits, or raise the cap at https://app.routerplus.com/console/billing; on cap breaches the exact reset time is in the message and x-tm-cap-reset
upstream_error (5xx)yesone retry after ~2 s; a failover already happened if one was possible
upstream_error (4xx)never blind-retrythe request itself was rejected — the message names the provider and its status; x-tm-upstream-status carries it
gateway_error (400)never blind-retrysend one output limit, between 1 and 32,768
gateway_error (500)noreport with the x-request-id
gateway_error (502, images)yes, oncenothing was billed; lower n or the quality
gateway_error (503)yeswait retry-after and resend — the request never reached a provider, so a retry cannot double-bill
gateway_error (504)yesresend; the deadline passed while candidates were being tried

Failure headers

HeaderWhenMeaning
x-request-idalwaysquote it in any report; also the id for GET /v1/generation?id=
x-tm-error-codeevery failurethe canonical class
x-tm-error-originlimit, provider and infrastructure failuresgateway_admission, upstream_quota, upstream, gateway_infrastructure or authorization; a request refused for its own shape (bad JSON, n>1, an unknown model) carries none
x-tm-limit-scope, x-tm-limit-kind, x-tm-limit-ida refused limitscope (platform, org, workspace, principal, key, model, pool, endpoint), the kind of limit (rpm, rps, rps_queue, tpm, concurrency, spend, health, …), and the limit's id when one exists
x-tm-attemptsafter ≥1 dispatchphysical provider attempts, failovers included
x-tm-upstream-statuswhen a provider answeredthe provider's own HTTP status
retry-after429 / 503seconds to wait — a rate limit says when its window frees; provider-sent values are honored but capped at 60; the all-cooling-down 503 sends 5; the ledger-journal 503 sends 2
x-tm-cap-resetspend-cap 429 onlyexact ISO instant the cap resets (first instant of the next UTC month)

x-tm-error-code is not CORS-exposed; browser code reads the class from the body.

Strict by design

Five refusals people trip on are deliberate — the alternative in each case is billing you for something wrong.

n>1 is a typed 400, not a degraded answer. The normalizer emits exactly one choice; accepting n: 4 and silently billing a garbled single-choice response would be dishonest.

bash
curl -s https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"claude-haiku-4-5","n":4,"messages":[{"role":"user","content":"hi"}]}'
json
{
  "error": {
    "code": "invalid_request",
    "message": "n>1 is not supported in v1; request a single choice",
    "type": "invalid_request_error",
    "metadata": { "error_type": "invalid_request" }
  },
  "request_id": "..."
}

Content a model does not take is a typed 400, not an answer that ignored it. A provider that does not read images can drop the image, answer the rest, and still bill the request. So an image, file, audio or video part goes only to a route whose listing takes that kind, and a part that no route of the model takes is refused before anything is reserved or sent. GET /v1/models lists what each model takes in architecture.input_modalities.

bash
curl -s https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"deepseek/deepseek-v4-flash","messages":[{"role":"user","content":[{"type":"text","text":"What is in this picture?"},{"type":"image_url","image_url":{"url":"https://example.com/cat.png"}}]}]}'
json
{
  "error": {
    "code": "invalid_request",
    "message": "messages[0].content[1]: model \"deepseek/deepseek-v4-flash\" does not take image input; it takes text",
    "type": "invalid_request_error",
    "metadata": { "error_type": "invalid_request" }
  },
  "request_id": "..."
}

No silent model substitution. An unknown model id is a 404 whose message echoes exactly the string you sent — we never quietly swap in a "close enough" model. The 404 is identical whether the id never existed or simply is not served, so the error is not an existence oracle.

A content-policy refusal never reroutes. A refusal is not shopped around to a more permissive provider — that is a hard routing rule with no exceptions. If you get content_policy, the fix is the prompt.

Images are all-or-nothing. A failed, over-cap or malformed generation is a 502 and bills nothing; n above 4 on the images route is a typed 400 like n > 1 on chat. See POST /v1/images/generations.

Failover, and what the error you see means

Before any output has reached you, the gateway fails over between deployments invisibly, at zero backoff. Classes that are failover-eligible: rate_limit, auth (upstream), model_unavailable, and 5xx upstream_error. Classes that never fail over: content_policy (never, as above), context_overflow (its own typed class — reshape the request instead), and provider 4xx rejections (the request is the problem). So if a failover-eligible class reaches you, every healthy candidate was tried — x-tm-attempts says how many.

Note

A relayed 402 (class auth) is never your fault. An upstream 401/402/403 means the marketplace's account with that provider has a credential or payment problem. Payability is treated as a health signal: the error is failover-eligible and counts against that provider's uptime, so an unpayable provider drains traffic automatically. You would only ever see it when every deployment serving the model has the same problem. On /v1/decisions most of these refusals come back as 503 model_unavailable instead, with metadata.retryable: false (see the table above).

Mid-stream failures (HTTP 200 already committed)

Once your answer has started streaming, the gateway never switches providers — a provider death mid-answer becomes exactly one terminal error event inside the 200 stream, and nothing follows it (no data: [DONE], no keep-alives, no more chunks).

OpenAI surface — a real chunk envelope with finish_reason: "error" and a top-level error, so SDKs raise and naive parsers still terminate cleanly:

json
{
  "id": "chatcmpl-...",
  "object": "chat.completion.chunk",
  "created": 1757000000,
  "model": "claude-haiku-4-5",
  "choices": [{ "index": 0, "delta": { "content": "" }, "finish_reason": "error" }],
  "error": {
    "code": "upstream_error",
    "message": "upstream failure",
    "type": "gateway_error",
    "metadata": { "error_type": "upstream_error" }
  }
}

Anthropic surface — a native error event, with the canonical class riding alongside the native type:

json
{ "type": "error", "error": { "type": "api_error", "message": "upstream failure", "error_type": "upstream_error" } }

On the OpenAI surface, any known usage and cost are emitted before the terminal error, so the wire stays billable-recomputable. The Anthropic wire has no legal usage slot outside message_delta; for an Anthropic-surface stream that died mid-answer, recompute from GET /v1/generation?id=<x-request-id>.

Retrying after a mid-stream error is your call — the answer never restarts on a different provider by itself.

Reading the class in code

python
import os
from openai import OpenAI, APIStatusError

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])

try:
    out = client.chat.completions.create(
        model="claude-haiku-4-5",
        messages=[{"role": "user", "content": "hello"}],
    )
except APIStatusError as e:
    err = e.response.json().get("error") or {}
    kind = (err.get("metadata") or {}).get("error_type") \
        or e.response.headers.get("x-tm-error-code")
    origin = e.response.headers.get("x-tm-error-origin")
    request_id = e.response.headers.get("x-request-id")
    # kind is one of the canonical classes in the table above

Retry discipline for agents

  • rate_limit → wait retry-after, retry once.
  • email_not_verified → do not loop; ask the human to open the verify link in their email.
  • insufficient_quota → do not loop; surface to a human (money). See Rate limits & spend caps.
  • usage_limited → do not loop; honor retry-after once, then tell the human to contact support.
  • usage_restricted → do not retry; tell the human to contact support.
  • upstream_error at 5xx → one retry after ~2 s, then surface with the x-request-id. At 4xx the request itself is the problem — never blind-retry.
  • upstream_error or gateway_error at 502 on the images route → one retry; nothing was billed.
  • gateway_error at 503 or 504 → wait retry-after (or ~2 s), retry; nothing was billed.
  • model_unavailable at 404 → pick a real id from GET /v1/models; at 502 one retry.
  • content_policy, context_overflow, invalid_request, gateway_error at 400 → never blind-retry; the request must change first.

A video that fails

A video job ends failed without an HTTP error: POST /v1/videos answered 200 when the job was accepted, and the failure arrives later, on GET /v1/videos/{id}, as status: "failed" with error: {code, message}. The code is one of video_generation_failed, content_policy_violation, video_cancelled, video_expired or video_timeout. A failed job costs nothing; content_policy_violation means change the prompt. See POST /v1/videos.

Markdown source for agents: /docs/errors.md · index at /llms.txt