Console
API reference/POST /v1/decisions

POST /v1/decisions

Typed answers from a decision model: Jev from TypeSafe, Mercury Decide from Inception, Bespoke Nimble v3 from Bespoke Labs, Clef and Clef-flash from Cloudflare, Decider 2B or Kev 4B from RouterPlus (our own), Perplexity Decider v1 27B from Perplexity, Sage from Levanto or GLiNER-2.5-Decide from Fastino, each in its own question format (eight of them share System One's).

/llms.txt

Ask a decision model typed questions about a state, and get one typed answer per question. A decision model does not write text. You send a state (a message, a ticket, a record) and your questions. Each answer is a choice, a score, a yes/no, an order, a set of tags or extracted fields, with the probability or confidence behind it. Use it for routing, classification, moderation and other decision points in your code.

The catalog has ten decision models:

ModelCatalog idMade byProvidersQuestion format
Jev 1.13typesafe/jev-1.13TypeSafetypesafe, then openrouter-decisionsSystem One questions: type is choice, score or noul
Mercury Decideinception/mercury-decideInceptionopenrouter-decisions-freeSystem One questions, exactly as for Jev
Bespoke Nimble v3bespokelabs/nimble-v3Bespoke LabsbespokelabsSystem One questions, as for Jev, within Nimble's own limits
Clefcloudflare/clefCloudflareworkers-aiSystem One questions, as for Jev, within Clef's own limits
Clef-flashcloudflare/clef-flashCloudflareworkers-aiSystem One questions, as for Jev, within Clef's own limits
Decider 2Brouterplus/decider-2bRouterPlusrouterplusSystem One questions, exactly as for Jev
Kev 4Brouterplus/kev-4bRouterPlusrouterplusSystem One questions, exactly as for Jev
Perplexity Decider v1 27Bperplexity/pplx-decider-v1-27bPerplexityperplexity-decisionsSystem One questions, as for Jev, within its own limits; it bills the state once per question
Sage 1.2levanto/sage-1.2LevantolevantoSystem One questions, as for Jev, within Sage's own limits; it bills the state once per question
GLiNER-2.5-Decidefastino/gliner-2.5-decideFastinofastinoA GLiNER schema: classifications, entities, structures, relations

The question format follows the model. Every call has the same envelope: model and state in, answers and usage out. The questions go in the model's own form: a questions map for the System One models, a schema object for GLiNER. What goes inside a question, and what comes back inside an answer, is the model's own format. Nine models share one format, System One (TypeSafe's): Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage. A body written for one runs on the others with only model changed, when it is inside each model's limits. A question in another model's format is a 400. The ten are different models, so the gateway never fails over from one to another.

The gateway calls these endpoints:

ModelOrderProviderEndpoint the gateway callsModel id sent upstream
Jev1TypeSafe (direct API)POST https://api.typesafe.ai/v1/systemonejev-1.13.0
Jev2OpenRouterPOST https://openrouter.ai/api/alpha/decisionstypesafe/jev-1.13-20260917
Mercury Decide1OpenRouterPOST https://openrouter.ai/api/alpha/decisionsinception/mercury-decide:free
Nimble1Bespoke Labs (direct API)POST https://api.bespokelabs.ai/v1/nimble/systemonenimble-v3
Clef1Cloudflare Workers AI (direct API)POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clefclef
Clef-flash1Cloudflare Workers AI (direct API)POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef-flashclef-flash
Decider 2B1RouterPlus (our own service, direct API)POST /v1/decisions on RouterPlus's endpointdecider-2b
Kev 4B1RouterPlus (our own service, direct API)POST /v1/decisions on RouterPlus's endpointkev-4b
Perplexity Decider1Perplexity (direct API)POST https://api.perplexity.ai/v1/decisionspplx-decider-v1-27b
Sage1Levanto (direct API)POST https://sage.levanto.ai/v1/systemonesage-latest
GLiNER1Fastino (direct API)POST https://api.fastino.ai/v1/chat/completionsfastino/GLiNER-2.5-Decide

For Jev you send one request shape. The gateway sends each provider the id it expects, and you get one answer shape back, whichever provider served it. If TypeSafe fails before it answers (overloaded, down, rate limited), the gateway tries OpenRouter. Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sage and GLiNER have one provider each, so their requests have no second provider to try. On Workers AI, {account_id} is our Cloudflare account, and the path names the model. Decider 2B and Kev 4B share one RouterPlus deployment: one pool and one health circuit serve both. The x-tm-provider header names the provider that answered.

To compare the models without code, use Decision in the playground. It takes one set of questions and translates it for each model.

Authentication

HeaderFormat
AuthorizationBearer tm_vk_...
x-api-keytm_vk_... — also accepted, same key

A missing or invalid key returns 401 with error_type: "auth".

Request

JSON body, at most 1 MB (413 request_too_large above that). The gateway checks every rule on this page before it reserves money or calls a provider. A refused request has no ledger row, no x-tm-attempts header and no charge.

ParameterTypeBehavior
modelstring, requiredA decision model id. A dated id of a listed decision model is accepted and routes as the model: typesafe/jev-1.13-20260917 (OpenRouter's dated id) as typesafe/jev-1.13, inception/mercury-decide-20260930 as inception/mercury-decide. OpenRouter's variant suffix (inception/mercury-decide:free) is not an id here: send inception/mercury-decide. Bespoke's own names (nimble-v3, nimble-latest) are not ids here either: send bespokelabs/nimble-v3. Nor are Cloudflare's (clef, clef-flash, @cf/cloudflare/clef, @cf/cloudflare/clef-flash): send cloudflare/clef or cloudflare/clef-flash. RouterPlus's own short names (decider-2b, kev-4b) are not ids here: send routerplus/decider-2b or routerplus/kev-4b. Perplexity's own name (pplx-decider-v1-27b) is not an id here either: send perplexity/pplx-decider-v1-27b. A chat, image or video model is a 400 that names its route. An unknown id is a 404.
staterequiredWhat the model evaluates. Its forms are the model's: see System One questions and GLiNER schema below.
questionsobject, the System One models, requiredA map of question id to question. You choose the ids; each answer comes back under the same id. At least one question. On GLiNER, questions is a 400 that says the model takes a GLiNER schema.
schemaobject, GLiNER only, requiredGLiNER's own schema. See GLiNER schema. On Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider or Sage, schema is a 400 that says the model takes questions.
threshold, include_confidence, include_spansGLiNER onlySee GLiNER options.
providerobjectRouting controls, read by the gateway and never forwarded: only, order, ignore, require_parameters and the rest — see Routing policies. The provider ids are typesafe and openrouter-decisions for Jev, openrouter-decisions-free for Mercury Decide, bespokelabs for Nimble, workers-ai for Clef and Clef-flash (not cloudflare, which is an OpenRouter host tag), routerplus for Decider 2B and Kev 4B, perplexity-decisions for Perplexity Decider (not perplexity, which is an OpenRouter host tag), levanto for Sage and fastino for GLiNER.

System One questions

Eight models take the same question format, TypeSafe's System One: Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B and Perplexity Decider. Everything in this section holds for each of the eight; only model differs, and Nimble, Clef, Clef-flash, the two RouterPlus models and Perplexity Decider have some limits of their own (see Bespoke Nimble v3, Clef and Clef-flash, Decider 2B and Kev 4B and Perplexity Decider v1 27B). The examples show Jev; put another System One model's id in model, for example inception/mercury-decide, bespokelabs/nimble-v3, cloudflare/clef, routerplus/decider-2b or perplexity/pplx-decider-v1-27b, and the same body runs on that model.

State. A non-empty string, or any JSON object or array (a chat log, a record, application state). Questions can point at a field by name: "Is `ticket.body` urgent?".

Questions. Every question has a type, instructions and, for most types, criteria. instructions is a non-empty string, or a JSON object or array: put the question in one field and the data it refers to in others.

typecriteriaThe answer
choiceRequired. An object of option name to description, 1 to 255 options. A description can be a string, JSON, or null when the name says enough.choice (the most likely option), probabilities (every option), confidence
scoreRequired. An ordered array of level descriptions, low to high, 2 to 10 levels.score (probability-weighted, can fall between levels), legend, probabilities, confidence
noulOptional. {"true": "...", "false": "..."} — what yes and no mean.noul: the probability that the answer is yes, 0 to 1

A question takes only type, instructions and criteria. Any other field in a question is a 400 that names it. A question with kind (the format of Levanto's own /decide API) is a 400 that says the model takes System One questions. Ask many questions in one call: the model reads the state once and answers every question against it, so one call with ten questions is cheaper and faster than ten calls. Perplexity Decider and Sage are the exceptions: they read and bill the state once per question, so on them one call saves round trips, not tokens (see Perplexity Decider v1 27B and Sage 1.2).

json
{
  "model": "typesafe/jev-1.13",
  "state": { "ticket": "Help! My payouts have been failing for 3 days." },
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle `ticket`?",
      "criteria": { "billing": "Payments, invoices, refunds", "technical": "Bugs, outages, integrations", "sales": null }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated is the customer?",
      "criteria": ["Calm", "Annoyed", "Angry"]
    },
    "urgent": {
      "type": "noul",
      "instructions": "Is this urgent?",
      "criteria": { "true": "Explicitly time-sensitive", "false": "No urgency expressed" }
    }
  }
}

Mercury Decide

Inception's Mercury Decide is a System One model: the state, the questions and the answers above are its format too. What is its own:

Mercury Decide
Catalog idinception/mercury-decide. The dated inception/mercury-decide-20260930 routes as the same model.
RouteOpenRouter's decisions endpoint only (POST https://openrouter.ai/api/alpha/decisions, provider id openrouter-decisions-free), as inception/mercury-decide:free. Inception's own API does not serve it yet.
Price$0 in and $0 out: OpenRouter serves only the model's free variant today, and the gateway passes that through. usage.cost is 0. See Billing.
Context32,768 tokens, the state and every question together.
Our limitThe deployment's pool takes 20 requests a minute in all, and each organization at most 15 of them; at most 5 requests in flight at once, 3 per organization; and 700,000 tokens a minute in all, 525,000 per organization. Tokens are held at the gateway's estimate of the body while a request runs (a state near the 32k context is held at about 44,000) and corrected to OpenRouter's count when it settles (about 33,000 for that state), so the 15-per-organization share of requests is the limit that binds, not tokens. A burst above any of these is a 429 rate_limit from the gateway, with retry-after and the x-tm-limit-* headers, before anything reaches OpenRouter: x-tm-limit-kind names the limit (rpm, tpm, concurrency, or cooldown for the pause after a relayed OpenRouter 429, next row) and x-tm-limit-id says whether it was the pool (pool:…) or your organization's share of it (pool-share:…).
OpenRouter's limitA free variant shares OpenRouter's daily cap on free requests across our whole account (1,000 a day). When that is used up, OpenRouter answers 429 until its daily reset, and the gateway relays it as 429 rate_limit with OpenRouter's retry-after and x-tm-upstream-status: 429. A relayed 429 pauses the deployment's pool for that retry-after, at least 1 s and at most 60 s: a request inside the pause gets the gateway's own 429 rate_limit, with x-tm-limit-kind: cooldown, x-tm-limit-id: pool:… and no x-tm-upstream-status. After the pause, requests reach OpenRouter again and meet the cap again, until its daily reset. A relayed 429 never counts against the deployment's health circuit, so the daily cap is 429s until the reset, not the 503 of a cooling-down deployment. The gateway does not count the daily cap itself, so a request can pass our pool and still meet it.
RetentionOpenRouter publishes no retention terms for this free endpoint, so the deployment declares prompt_logging: "retained". See Data policy.

A free endpoint is, in OpenRouter's own words, not production-suitable. For a decision that must come back, keep Jev as your System One fallback in your own code: the gateway never fails over between models.

Bespoke Nimble v3

Bespoke Labs' Nimble is a System One model: the state, the questions and the answers above are its format too. The gateway calls Bespoke's own API. What is its own:

Bespoke Nimble v3
Catalog idbespokelabs/nimble-v3. Bespoke's names nimble-v3 and nimble-latest are not ids here.
RouteBespoke Labs' API only (POST https://api.bespokelabs.ai/v1/nimble/systemone, provider id bespokelabs), as nimble-v3.
Price$0.04 per million input tokens. Output is free. This is Bespoke's own price, with no markup. See Billing.
Questions1 to 64 questions per call. A choice takes 2 to 255 options: a choice with one option, which Jev takes, is refused. A score takes 2 to 10 levels through the gateway, as for every System One model (Nimble itself takes up to 255). The gateway does not check Nimble's own limits: a body above them is Bespoke's 422, relayed as an upstream_error, and costs nothing. (The Playground's Decision mode does check them, before a run.)
Structured contentinstructions and a choice option's description may be a JSON object or array, as on Jev. Bespoke's published schema types them as strings, but Nimble takes them: a probe with both answered 200 (2026-10-02).
Context32,768 tokens for each question's prompt, the state included. Nimble never cuts a prompt: a longer one is refused. Bespoke counts the state and the questions once per call, not once per question.
BusyBespoke answers 503 with Retry-After while Nimble starts, 529 when it is busy (Bespoke says: retry after about one second) and 502 when the model fails. Each is a retryable upstream_error that counts against the deployment's health circuit, like every 5xx: two in a row open it for 30 s, and in that window every request but one probe gets 503 gateway_error, "all deployments cooling down". Retry after a second or two, with backoff. The gateway does not relay Bespoke's Retry-After on a 503. See Errors.
NumbersFull precision: Jev's numbers have 2 decimals, Nimble's are not rounded. A very small probability can come back in exponent form, for example 3.7e-06. confidence is 1 when one option has all the probability and 0 when every option is equally likely. It is not the chance that the answer is right.
Our limitThe deployment's pool takes 7 requests in flight at once, and each organization at most 5 of them (Bespoke allows our account 8; one is kept for our test environment). It also takes 600 requests a minute (450 per organization) and 15,000,000 tokens a minute (11,250,000 per organization). A burst above any of these is a 429 rate_limit from the gateway, with retry-after and the x-tm-limit-* headers, before anything reaches Bespoke; x-tm-limit-kind names the limit (rpm, tpm, concurrency, or cooldown for the pause after a relayed Bespoke 429, next row).
Bespoke's limitBespoke runs at most 8 requests at once for our whole account. Above that it answers 429 with Retry-After. The gateway relays it as 429 rate_limit with x-tm-upstream-status: 429, and pauses the deployment's pool for that retry-after, at least 1 s and at most 60 s. A relayed 429 never counts against the deployment's health circuit.
RetentionBespoke keeps request content for up to 30 days, to debug the service and check for abuse, and then deletes it. The deployment declares prompt_logging: "retained". Bespoke does not train on API content unless an organization opts in; ours does not. See Data policy.

Clef and Clef-flash

Cloudflare's Clef and Clef-flash are System One models: the state, the questions and the answers above are their format too. Clef-flash is the smaller model (9 billion parameters, against Clef's 27 billion): it is faster and costs less. The gateway calls Cloudflare's own API, Workers AI. What is their own:

Clef and Clef-flash
Catalog idscloudflare/clef and cloudflare/clef-flash. Cloudflare's names clef, clef-flash, @cf/cloudflare/clef and @cf/cloudflare/clef-flash are not ids here.
RouteCloudflare Workers AI's REST API only, provider id workers-ai: POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef as clef, and .../@cf/cloudflare/clef-flash as clef-flash. {account_id} is our Cloudflare account. The path picks the model, and the body's model must be the same model; the gateway sets both. Cloudflare puts each answer inside its own envelope (result, success, errors, messages). The gateway opens it, so you get Jev's answer shape. A 200 whose envelope does not say success: true is a 502 upstream_error and costs nothing.
PriceClef: $0.24 per million input tokens. Clef-flash: $0.09 per million input tokens. Output is free: Cloudflare reports output_tokens: 0 on every answer. This is Cloudflare's own price, with no markup. See Billing.
Questions1 to 64 questions per call. Each question id (a key of questions) is 1 to 100 letters, digits, _, . or -: it must match ^[A-Za-z0-9_.-]{1,100}$. A choice takes 2 to 255 options: a choice with one option, which Jev takes, is refused. A score takes 2 to 10 levels. The gateway does not check Clef's own limits: a body outside them is Cloudflare's 422 (or 400), relayed as an upstream_error with Cloudflare's words, and costs nothing. (The Playground's Decision mode does check them, before a run.)
Structured contentinstructions, a score level and a choice option's description may be a JSON object or array, as on Jev, and an option's description may be null. Cloudflare's schema says so. A probe with a JSON object as instructions and JSON option descriptions answered 200 (2026-10-02).
Context65,536 tokens, the state and every question together. Before the model runs, Cloudflare estimates the request's tokens at about 4 characters a token. It refuses a request estimated above 65,536 tokens (about 262,000 characters) with a 413, which the gateway returns as 413 context_overflow; it costs nothing. Cloudflare's schema says that a long state is cut to fit, but in our test (2026-10-02) a request of 522,286 characters was refused, not cut.
ImagesNot accepted through the gateway. Clef's images field, Cloudflare's addition to System One, is dropped and recorded in x-tm-dropped-params, never forwarded. With "provider": {"require_parameters": true} it is a 400.
request_idDo not send a top-level request_id. The gateway drops and records it, and never forwards it: Workers AI reads that field as a lookup in its queue of async requests, and answers 404.
Numbers4 decimals: Jev's numbers have 2, and Nimble's are not rounded. confidence is on Clef's own scale, not on Jev's. Cloudflare says only that it is derived from the probabilities. On one live answer (a choice of three options, the top one at 0.8274), Clef's confidence was 0.5654, where Jev's formula gives 0.7411. So a threshold that you tuned on Jev's confidence does not transfer to Clef: tune it on Clef, or use the probabilities. confidence is not the chance that the answer is right.
BusyWhen Workers AI is busy, Cloudflare answers 429 with its code 3040, "Capacity temporarily exceeded, please try again." The gateway relays it as 429 rate_limit, with Cloudflare's words and x-tm-upstream-status: 429, and the deployment's pool pauses for Cloudflare's Retry-After, at least 1 s and at most 60 s. Without a Retry-After, the relayed 429 has no retry-after either, and the pool pauses for 1 s. A relayed 429 never counts against the deployment's health circuit. Retry after a second or two, with backoff. A 5xx (for example Cloudflare's 500, "Model execution failed") is a retryable upstream_error that counts against the circuit, like every 5xx: two in a row open it for 30 s. A 408, Cloudflare's own timeout, is a retryable 504 upstream_error that does not count. See Errors.
Our limitClef and Clef-flash share one pool. It takes 200 requests a minute (150 per organization), at most 8 requests at a time (6 per organization) and 2,000,000 tokens a minute (1,500,000 per organization). A burst above any of these is a 429 rate_limit from the gateway, with retry-after and the x-tm-limit-* headers, before anything reaches Cloudflare; x-tm-limit-kind names the limit (rpm, tpm, concurrency, or cooldown for the pause after a relayed Cloudflare 429).
Cloudflare's limitCloudflare publishes 300 requests a minute for Clef's task type, Text Generation. It does not say whether Clef and Clef-flash share that limit, so our pool stays below it. Cloudflare's 429 with code 3036 says that our account's free daily allocation is used up. The gateway answers it with 503 model_unavailable and metadata.retryable: false: that is our account, never yours. It never counts against the deployment's health circuit.
RetentionCloudflare says that it does not read, store or train on requests to Clef, or on their responses. The deployment declares prompt_logging: "none". The gateway calls Workers AI directly, never through Cloudflare's AI Gateway, which logs prompts and responses by default. See Data policy.

Decider 2B and Kev 4B

RouterPlus is our own decision-model service. Its two models run on our own GPUs in the United States, and both are System One models: the state, the questions and the answers above are their format too. Each reads English. What is their own:

Decider 2BKev 4B
Catalog idrouterplus/decider-2brouterplus/kev-4b
Name at RouterPlusdecider-2bkev-4b
Input, per million tokens$0$0
Context25,600 tokens; a longer input is cut (see below)8,192 tokens; a longer input is refused

What holds for both:

Decider 2B and Kev 4B
RouteRouterPlus's endpoint only (POST /v1/decisions there, provider id routerplus). RouterPlus's short names in the table above are not ids here. One deployment serves both models, so one pool, one health circuit and one pause after a 429 cover both.
PriceFree: input and output are $0, and usage.cost is 0 on every call. RouterPlus is our own service, so the price is ours to set, not a provider's price passed through. See Billing.
QuestionsThe gateway checks the System One rules above. RouterPlus takes every form that Jev takes, on both models: instructions as a string or as a JSON object or array, score levels as strings or as JSON objects (the answer's legend gives the levels as you sent them), and a choice option's description as a string, JSON or null. In probes on 2026-10-02, each model also answered a choice with one option, a choice of 255 options, a score of 2 and of 10 levels, a yes/no with described true and false, a JSON object or array as the state, and 200 questions in one request. A question that RouterPlus calls malformed is its 400, relayed with RouterPlus's reason for each question id (see Errors); it costs nothing.
Long inputDecider 2B reads up to 25,600 tokens and cuts a longer input: it keeps the question and its options first, then the start of the state. It drops the rest and does not return an error. Kev 4B never cuts: above 8,192 tokens it refuses the request with a 400 (see Errors). The gateway bills RouterPlus's usage.input_tokens, and RouterPlus counts there only the tokens the model can read: at most the model's context. In a test on 2026-10-02, a Decider 2B request of 30,053 tokens billed 25,600. RouterPlus can refuse a long input in place of cutting it (its truncate field), but the gateway drops that field (see Parameters that do not apply), so through the gateway a long input on Decider 2B is always cut. Put what matters at the start of the state, or use Kev 4B for an input up to 8,192 tokens and Decider 2B up to 25,600.
Response cacheRouterPlus keeps each answer, with its usage, for 600 s (10 minutes). The cache key is the RouterPlus API key, the model and the exact request body, byte for byte. The gateway sends every request with our one RouterPlus key, so an identical request from any of our buyers within 600 s can get the cached answer: the same answers, and the same input count in usage. Your response keeps its own id. RouterPlus marks it with usage.cached: true and the header X-Cache: HIT in its own response. A cached answer costs $0: the listing prices cached input at $0, and the gateway bills a cached answer's input tokens at that price. The gateway's response marks it with usage.cached_input_tokens, equal to usage.input_tokens. RouterPlus skips its cache for a request with "cache": false or a Cache-Control: no-cache header. The gateway forwards neither, so you cannot skip the cache through the gateway.
Request idRouterPlus sends X-Request-Id (req_…) on every answer. The gateway keeps it with the request's record as the provider's id. Your response's id and its x-request-id header are the gateway's own.
Cold startAfter a quiet period, the first request starts the model on a GPU. This takes 15 to 20 s. The gateway waits up to 60 s for a decision, so a cold start is a slow 200, not an error.
Our limitThe deployment's pool takes 11,400 requests a minute, 15,000,000 tokens a minute and 59 requests in flight at once. Each organization may use 75 % of each: 8,550 requests a minute, 11,250,000 tokens a minute, 44 at once. The 59 in flight is the burst limit: a pool has no limit per second, and 59 requests in flight at RouterPlus's typical 0.31 s a call carry about 190 a second, just above the pool's 11,400 a minute; with our test environment's 10 a second, that stays within RouterPlus's own limit (next row). Your organization's own limits bind first: by default 300 requests a minute, 1,000,000 tokens a minute and 8 in flight (see Rate limits & spend caps). A burst above any of these is a 429 rate_limit from the gateway, with retry-after and the x-tm-limit-* headers, before anything reaches RouterPlus; x-tm-limit-kind names the limit (rpm, tpm, concurrency, or cooldown for the pause after a relayed RouterPlus 429, next row).
RouterPlus's limitAbout 200 requests a second for the whole endpoint, shared with our test environment. A cached answer comes back sooner than a computed one, so a burst of repeated requests can still reach it. Above it RouterPlus answers 429. The gateway relays it as 429 rate_limit with x-tm-upstream-status: 429, and pauses the deployment's pool for that retry-after, at least 1 s and at most 60 s. A relayed 429 never counts against the deployment's health circuit.
Busy or failedA 5xx from RouterPlus is a retryable upstream_error that counts against the deployment's health circuit: two in a row open it for 30 s for both models, and in that window every request but one probe gets 503 gateway_error, "all deployments cooling down". Retry after a few seconds, with backoff. A RouterPlus 401, 402 or 403 refuses our key, never yours: you get 503 model_unavailable, as from every System One provider. See Errors.
RetentionRouterPlus writes no request logs. Its answer cache keeps each answer and its usage for 600 s, under a key made from our RouterPlus key and the request. The cache is all it keeps, so the deployment declares prompt_logging: "none". Nothing is used for training. See Data policy.

Perplexity Decider v1 27B

Perplexity's Decider v1 27B is a System One model: the state, the questions and the answers above are its format too. The gateway calls Perplexity's own API. One thing sets it apart from the rest of the family, and it changes the cost: it reads and bills the state once per question. What is its own:

Perplexity Decider v1 27B
Catalog idperplexity/pplx-decider-v1-27b. Perplexity's own name pplx-decider-v1-27b is not an id here.
RoutePerplexity's API only (POST https://api.perplexity.ai/v1/decisions, provider id perplexity-decisions), as pplx-decider-v1-27b. The provider id is not perplexity, which is an OpenRouter host tag. Perplexity answers in Jev's shape, with no envelope. The gateway keeps Perplexity's x-request-id with the request's record as the provider's id.
Price$0.04 per million input tokens. Output is free. This is Perplexity's own price, with no markup. See Billing.
The state, once per questionPerplexity runs one prompt for each question, and each prompt holds the whole state. It bills the state in each one. In a test on 2026-10-02, a state of 1,843 tokens billed 1,843 tokens with one question and 9,215 with five. So on this model one call with ten questions saves round trips, not tokens: it costs about what ten calls cost. The gateway holds the state once per question too (see Our limit below). The question ids are not billed. output_tokens is one per question, at $0. A choice with one option runs no prompt and costs nothing: a request of only such questions comes back with 0 tokens and costs $0.
Questions1 to 128 questions per call. Above 128 is Perplexity's 400, "Each request needs between 1 and 128 questions", relayed as an upstream_error with Perplexity's words; it costs nothing. A choice takes 1 to 255 options, as on Jev. A score takes 2 to 10 levels. A question id is any key the gateway takes. The gateway does not check the 128 (the Playground's Decision mode does, before a run).
Structured contentinstructions, a score level and a choice option's description may be a JSON object or array, as on Jev, and an option's description may be null. A probe with a JSON object as instructions, JSON option descriptions and JSON score levels answered 200 (2026-10-02).
Context262,144 tokens for each question's prompt: the state and that one question. A request with several questions can bill more than that in all: a state of 52,348 tokens with six questions billed 314,088 tokens (2026-10-02). Above the limit, Perplexity refuses the request with a 400, "Input length (262144) exceeds or equals model's maximum context length (262144)". The gateway returns it as 400 context_overflow, with metadata.retryable: false; it costs nothing. Perplexity does not cut a long state.
Image partsDecision models take text and JSON. An object with "type": "image_url", anywhere in the state or in a question, is a 400 from the gateway before any money is reserved, for example state[1]: an object with "type": "image_url" is an image part, and model "perplexity/pplx-decider-v1-27b" takes text and JSON only. Other JSON, a field named type with another value included, goes as you sent it.
NumbersFull precision, as Nimble's. On a choice, confidence is Jev's: (p_max − 1/n) / (1 − 1/n), so a threshold that you tuned on Jev's choice confidence transfers. On a score it does not. Perplexity's score confidence is `1 − Σ pᵢ·i − m/ ((1/L)·Σⱼj − m), where mis the most likely level andLthe number of levels; Jev measures the second sum about the central level, not aboutm. The two agree only when the most likely level is the central one. On one live answer (three levels at 0.756, 0.193 and 0.051), Perplexity's confidencewas 0.7051, where Jev's formula gives 0.5577. So tune a score threshold on Perplexity Decider, or use the probabilities.confidence` is not the chance that the answer is right.
BusyPerplexity answers 504 when the model does not answer in about a minute; the body is an HTML page. The gateway returns a retryable 504 upstream_error that counts against the deployment's health circuit, like every 5xx: two in a row open it for 30 s, and in that window every request but one probe gets 503 gateway_error, "all deployments cooling down". The gateway's own wait for a decision (60 s) also ends in a 504. In Perplexity's tests a few hundred input tokens answered in under 2 s, and a prompt near the context in 23 s. A Perplexity 401, 402 or 403 refuses our key, never yours: you get 503 model_unavailable, as from every System One provider. See Errors.
Our limitThe deployment's pool takes 540 requests a minute, at most 8 requests at a time and 15,000,000 tokens a minute. Each organization may use 75 % of each: 405 requests a minute, 6 at a time, 11,250,000 tokens a minute. Your organization's own limits bind first: by default 300 requests a minute, 1,000,000 tokens a minute and 8 in flight (see Rate limits & spend caps). Tokens are held at the gateway's estimate while a request runs, with the state counted once per question, and corrected to Perplexity's count when it settles. So a long state with many questions can be larger than a token limit on its own: such a request is a 429 rate_limit with x-tm-limit-kind: tpm before anything reaches Perplexity, and it cannot pass as it is. Ask fewer questions per call, or send a shorter state. A burst above any limit is a 429 rate_limit from the gateway, with retry-after and the x-tm-limit-* headers; x-tm-limit-kind names the limit (rpm, tpm, concurrency, or cooldown for the pause after a relayed Perplexity 429, next row).
Perplexity's limitPerplexity allows our organization 10 requests a second, over a rolling second, shared with our test environment. Our pools keep to that as a minute budget (600 a minute in all), but they do not count seconds: a burst inside one second can pass 10, and then Perplexity's 429 limits it. It also limits large bursts of tokens, and publishes no number for that. Above a limit Perplexity answers 429, "Request rate limit exceeded, please try again later.", with Retry-After: 1. The gateway relays it as 429 rate_limit with x-tm-upstream-status: 429, and pauses the deployment's pool for that retry-after, at least 1 s and at most 60 s. A relayed 429 never counts against the deployment's health circuit.
RetentionPerplexity says that it does not retain query data sent through its API and does not train on it: its API has zero day retention of prompt data by default. Its compute runs on AWS in North America. The deployment declares prompt_logging: "none". See Data policy.

Sage 1.2

Levanto's Sage is a System One model: the state, the questions and the answers above are its format too. The gateway calls Levanto's System One API. Like Perplexity Decider, it reads and bills the state once per question. What is its own:

Sage 1.2
Catalog idlevanto/sage-1.2. Levanto's own names (sage-latest, levanto-sage-v1.3) are not ids here.
RouteLevanto's API only (POST https://sage.levanto.ai/v1/systemone, provider id levanto), as sage-latest. Levanto answers in Jev's shape.
Price$0.05 per million input tokens. Output is free. This is Levanto's own price, with no markup. See Billing.
The state, once per questionLevanto reads the state once for each question and bills it each time. So on this model one call with ten questions saves round trips, not tokens. The gateway holds the state once per question too.
QuestionsA choice takes 2 to 120 options and a score 2 to 26 levels. Every question gets an answer, never null. Levanto refuses a request outside its limits; the gateway relays the refusal with Levanto's words, and it costs nothing.
Not offeredLevanto's own /decide format (kind, sort, tags), images, grounding and reasoning. A question with kind is a 400; reasoning and latency_mode are dropped and recorded.
RetentionThe deployment declares prompt_logging: "retained". See Data policy.

GLiNER schema

GLiNER-2.5-Decide is a small encoder model. It reads one text and fills in a schema: it classifies the text against your labels, and it can extract entities, structured fields and relations from it.

State. Text only: a non-empty string, or {"kind": "text", "value": "..."}. A list, an image or any other state is a 400 that names the field. To judge a record, send it as text, for example the JSON as a string.

Schema. GLiNER takes no questions. It takes schema, Fastino's own schema object. The gateway checks it and forwards it as it is. schema has only these four keys; send at least one, and each key you send must be non-empty (an empty list or object is a 400 that names the key). A flat list in place of the object is a 400 (Fastino has deprecated it).

KeyFormLimits
classificationsA list of {"task", "labels", "multi_label"?, "top_k"?, "cls_threshold"?}1 to 50 tasks. task is 1 to 256 characters, unique in the schema, and not __proto__. labels is 1 to 100 unique non-empty strings. multi_label is a boolean. top_k is an integer, 1 or more. cls_threshold is a number from 0 to 1. Any other key is a 400 that names it.
entitiesA list of entity names, or of {"name", "description"?}1 to 50 entities. Each name is a non-empty string, unique in the list.
structuresAn object of structure name to a list of fields, each "field::type::description"1 to 50 structures. A name is 1 to 256 characters and not __proto__. 1 to 50 non-empty field strings per structure.
relationsA list of relation names, or an object of name to {"description"?, "threshold"?}1 to 50 unique non-empty names. threshold is a number from 0 to 1. A head or tail key is a 400.

Labels are plain strings. A label cannot be an object with a description. To describe a label, put the description in the label itself: "credit: store credit". Descriptions matter: on the same refund ticket, the labels refund, credit, deny chose refund, and the same labels with descriptions chose credit. The answer carries the label text exactly as you sent it.

The task name is the question. A classification has no instructions field. GLiNER reads the task name, so a task name can be a question: a task named "Does the customer ask for money back?" with the labels ["yes", "no"] is a yes/no question.

Result names share one namespace. GLiNER puts every result at the top level of its answer: each task under its task name, each structure under its structure name, the entities under entities and the relations under relation_extraction. Two results with one name would overwrite each other at Fastino without an error. So the gateway refuses a collision with a 400 that names both, for example a task named entities in a request that also asks for entities.

json
{
  "model": "fastino/gliner-2.5-decide",
  "state": "I was charged twice for my March invoice. Please refund the order from Acme Corp. Signed, Jane Doe.",
  "schema": {
    "classifications": [
      { "task": "action", "labels": ["refund", "credit", "deny"] },
      { "task": "issues", "labels": ["double_charge", "angry_customer", "fraud"], "multi_label": true }
    ],
    "entities": ["person", { "name": "organization", "description": "business or institution name" }]
  }
}

GLiNER options

Three optional top-level fields go to Fastino as they are. Leave them out to use Fastino's defaults.

FieldValuesWhat it does
thresholdA number from 0 to 1. Default 0.5.The confidence a result needs to be returned. A task's own cls_threshold wins for that task. Lower values return more results; higher values return fewer, surer ones.
include_confidenceBoolean. Default true.With false, a single-label answer is the bare label string, and a multi-label answer is a list of label strings.
include_spansBoolean. Default true.With true, each entity carries start and end: its character offsets in the text.

A single-label task returns its top label only; top_k does not add more labels to the answer. When no label reaches the threshold, the task's answer is null: set "cls_threshold": 0 on the task to always get the top label and its confidence. A task with multi_label: true returns every label at or above the threshold, so "cls_threshold": 0 returns every label with its confidence.

Store

Fastino keeps each inference unless the request says store: false. The gateway always sends store: false. A store field that you send is dropped and recorded, never forwarded. Our Fastino account also has Zero Data Retention on, so Fastino does not train on your text. See Data policy.

Not offered yet

These parts of Fastino's API are not reachable through the gateway today:

  • Fine-tuned GLiNER models (Fastino's training-job ids).
  • Fastino's own batch and async routes.
  • Several messages. The gateway sends Fastino one user message: the state. GLiNER does not read a system message.

Parameters that do not apply

Decision models take no sampling or output controls. The gateway forwards no temperature, max_tokens, stream, seed, user or metadata to them, and none offers prompt caching, so cache_control has nothing to act on. RouterPlus's answer cache (see Decider 2B and Kev 4B) is not prompt caching: it works by itself on a whole identical request, and cache_control does not reach it. A top-level field that the model does not take is dropped and recorded (D8 §2), never forwarded. RouterPlus's own truncate and cache fields are dropped in this way:

ModelTop-level fields it takesExamples of fields dropped
Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sagemodel, state, questions, providertemperature, reasoning, latency_mode, images, request_id, truncate, cache
GLiNERmodel, state, schema, provider, threshold, include_confidence, include_spansstore, temperature, messages

The dropped names are in the x-tm-dropped-params response header and on the attempt's ledger row. To refuse instead, send "provider": {"require_parameters": true}: a would-be drop is then a 400 before any money is reserved.

Response

HTTP 200, application/json. Every decision model uses the same envelope: id, object, model, answers and usage.

System One answers

The shape is the same from both of Jev's providers, from Mercury Decide (with "model": "inception/mercury-decide" and "cost": 0), from Nimble (with "model": "bespokelabs/nimble-v3" and its numbers at full precision), from Clef and Clef-flash (with "model": "cloudflare/clef" or "model": "cloudflare/clef-flash", its numbers to 4 decimals and "output_tokens": 0), from Decider 2B and Kev 4B (with their catalog ids in model) and from Perplexity Decider (with "model": "perplexity/pplx-decider-v1-27b", its numbers at full precision and one output token per question):

json
{
  "id": "5d0c1c4e-3f7a-4c55-9d1b-2f0e7f6f2a10",
  "object": "decision",
  "model": "typesafe/jev-1.13",
  "answers": {
    "team": { "type": "choice", "choice": "billing",
              "probabilities": { "billing": 0.88, "technical": 0.12, "sales": 0 }, "confidence": 0.81 },
    "frustration": { "type": "score", "score": 1.05,
                     "legend": { "0": "Calm", "1": "Annoyed", "2": "Angry" },
                     "probabilities": { "0": 0, "1": 0.95, "2": 0.05 }, "confidence": 0.92 },
    "urgent": { "type": "noul", "noul": 0.95 }
  },
  "usage": { "input_tokens": 436, "output_tokens": 71, "cost": 0.000018 }
}

answers holds one typed answer per question, under the ids you sent, exactly as the provider returned them.

GLiNER answers

answers is GLiNER's own result, verbatim: the JSON that Fastino returns as a string in choices[0].message.content, parsed into an object. The gateway adds nothing and removes nothing. For the request in GLiNER schema:

json
{
  "id": "0b7e4c2a-6f1d-4a39-9c85-3e2d1f7a8b64",
  "object": "decision",
  "model": "fastino/gliner-2.5-decide",
  "answers": {
    "action": { "label": "refund", "confidence": 0.9327635765075684 },
    "issues": [ { "label": "double_charge", "confidence": 0.872347354888916 } ],
    "entities": {
      "person": [ { "text": "Jane Doe", "confidence": 0.99609375, "start": 90, "end": 98 } ],
      "organization": [ { "text": "Acme Corp", "confidence": 0.98828125, "start": 71, "end": 80 } ]
    }
  },
  "usage": { "input_tokens": 33, "output_tokens": 91, "cost": 0 }
}
You askedKey in answersValue
A single-label taskThe task name{"label", "confidence"}: the top label. null when no label reaches the threshold (0.5 unless you set one; "cls_threshold": 0 always returns the top label). With include_confidence: false, the bare label string.
A multi-label taskThe task name[{"label", "confidence"}, ...]: every label at or above the threshold, possibly none. With include_confidence: false, a list of label strings.
EntitiesentitiesAn object of entity name to [{"text", "confidence", "start", "end"}, ...]. start and end are there when include_spans is on.
A structureThe structure nameA list of records, each field as {"text", "confidence"}, for example {"order": [{"id": {"text": "4411", "confidence": 1.0}}]}.
Relationsrelation_extraction{"relations": [...]}.
  • A confidence is not a calibrated probability. It is GLiNER's score for that label or span, from 0 to 1. Multi-label confidences are independent and do not sum to 1. Compare them within one model, not with System One's probabilities.
  • GLiNER has no "not sure" answer. A low confidence is the sign that it is unsure.
  • The gateway accepts a 200 only when the content parses to a JSON object with every task and structure name you asked, plus entities and relation_extraction when you asked for them. Anything else is an upstream_error and costs nothing.

Usage and headers

FieldMeaning
idThe gateway's request id, also in the x-request-id header. Look the request up with GET /v1/generation?id=.
modelThe catalog id that was billed. The provider's own name for the model is in the x-tm-upstream-model header.
usage.input_tokensInput tokens, billed at the model's input rate. For the System One models and GLiNER, the provider's count (Fastino's prompt_tokens); on Decider 2B, at most the model's context; on Perplexity Decider and Sage, the state once per question.
usage.output_tokensOutput tokens, billed at the model's output rate. Every decision model prices output at $0; Sage, Clef and Clef-flash report 0, and Perplexity Decider one per question. For GLiNER, Fastino's completion_tokens.
usage.cached_input_tokensPresent only on an answer from RouterPlus's cache (Decider 2B, Kev 4B). The part of usage.input_tokens that the cache answered: all of it. These tokens are billed at the cached-input price, $0.
usage.costUSD, the full debit for the request, after any discount, including any attempt that failed over before this one. 0 on BYOK. Absent on static dev keys.
usage.cost_before_discountPresent only when a discount applied. What the request costs at the list price.
usage.discount_percentPresent only when a discount applied. The percent taken off.

Response headers: x-request-id, x-tm-provider (typesafe, openrouter-decisions, openrouter-decisions-free, bespokelabs, workers-ai, routerplus, perplexity-decisions, levanto or fastino), x-tm-served-by (the provider OpenRouter names, when it names one: Inception for Mercury Decide), x-tm-upstream-model (the provider's own name for the model, for example clef or clef-flash from workers-ai, decider-2b from routerplus, pplx-decider-v1-27b from perplexity-decisions), x-tm-attempts, x-tm-upstream-status, x-tm-discount-percent when a discount applied and, when something was dropped, x-tm-dropped-params.

Billing

A decision is metered like any other request: key limits, the reservation, the ledger and usage.cost all work as on the chat routes. Every decision model is priced per million tokens. Jev, Nimble, Clef, Clef-flash, Perplexity Decider, Sage and GLiNER charge for input only. Mercury Decide is $0 both ways while OpenRouter serves only its free variant; a paid variant, or Inception serving it directly, would be a new listed price, and this page would say so. Decider 2B and Kev 4B run on RouterPlus, our own service, and they are free: $0 in and out, so usage.cost is 0 on every call. Each charge rounds down to the micro-dollar. Perplexity Decider bills the state once per question (see Perplexity Decider v1 27B): its cost grows with the number of questions as well as with the state.

ModelInput, per million tokensOutput, per million tokensWhat one call costs
Jev$0.042$0A three-question call of about 400 input tokens costs about $0.000017.
Mercury Decide$0$0Nothing. The call is still metered, limited and recorded in GET /v1/generation, and it still needs an organization with credit.
Nimble$0.04$0A three-question call of about 340 input tokens costs $0.000013. Bespoke's own price, with no markup.
Clef$0.24$0A call of 400 input tokens costs $0.000096. Cloudflare's own price, with no markup.
Clef-flash$0.09$0A call of 400 input tokens costs $0.000036. Cloudflare's own price, with no markup.
Decider 2B$0$0Nothing, as on Mercury Decide.
Kev 4B$0$0Nothing, as on Mercury Decide.
Perplexity Decider$0.04$0Five questions about a state of 1,843 tokens bill 9,215 input tokens: $0.000368. Perplexity's own price, with no markup.
Sage$0.05$0Three questions about a short ticket read 204 input tokens: $0.00001. Levanto's own price, with no markup.
GLiNER$0.03$0A call of 1,000 input tokens costs $0.00003. A call of 33 input tokens, as in the example in GLiNER answers, rounds down to $0.

If a provider answers 200 without a usage object, the gateway bills its own estimate, the same amount it reserved, and marks the attempt estimated. For GLiNER it comes from the size of the text and from the tasks, labels and extraction keys in the schema. For Perplexity Decider and Sage it counts the state once per question. A 200 that does not answer every question, or whose GLiNER content is not the object described in GLiNER answers, is an upstream_error and costs nothing. After a provider answers 200, the gateway never sends the same request to a second provider.

Discount

A discount may apply to decision models: to every decision model, or to some of them, each at its own percent. It applies from every key, including the playground's. The list price does not change, and the discount can change or end: read it from each response, not from this page. A discounted model's row in /api/models.json carries discount_percent, and its page shows the percent beside the struck list price.

  • usage.cost is what you pay, after the discount.
  • usage.cost_before_discount is the list price of the request.
  • usage.discount_percent and the x-tm-discount-percent header give the percent.
  • When no discount applies, these three are absent.

A discounted charge is the list charge × (100 − percent) / 100, rounded down to the micro-dollar. At 100 % it is 0, and the request is still metered, limited and recorded in GET /v1/generation. The discount is for organizations with credit: a balance of $0 is refused with 429 insufficient_quota whatever the discount, so a new account must complete verified browser sign-in and buy credits before its first decision.

On Mercury Decide, Decider 2B and Kev 4B the list price is already $0, so the discount changes nothing: usage.cost and usage.cost_before_discount are both 0. When a discount covers one of them, the percent is still reported, as on every model it covers, so your code can read one shape for all ten.

Errors

Errors use the gateway's normal shape — see Errors.

Statuserror.codeWhen
400invalid_requestThe body failed a rule above: a question in another model's format, a field a question does not take. The message names the field, for example questions.team.criteria: at most 255 options.
400invalid_requestThe request does not fit the model's grammar: schema on Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider or Sage (the message says the model takes questions), or questions on GLiNER (the message says the model takes a GLiNER schema).
400invalid_requestGLiNER only: a state that is not text, a flat-list schema, a schema key or classification key GLiNER does not take, a limit above, two tasks with one name, or two results with one name (the message names both).
400invalid_requestThe model is not a decision model. The message names the route that serves it.
400invalid_requestPerplexity Decider: an object with "type": "image_url" anywhere in the state or in a question. Decision models take text and JSON. The message names where it is, for example state[1]: an object with "type": "image_url" is an image part, and model "perplexity/pplx-decider-v1-27b" takes text and JSON only. It is refused before any money is reserved.
400context_overflowPerplexity Decider: one question's prompt (the state and that question) is above 262,144 tokens. Perplexity's 400, relayed with its words, for example upstream perplexity-decisions returned 400: Input length (262144) exceeds or equals model's maximum context length (262144). metadata.retryable is false: make the state shorter. It costs nothing.
404model_unavailableNo decision model has this id.
413request_too_largeThe body is larger than 1 MB.
413context_overflowClef and Clef-flash: Cloudflare estimated the request above the model's 65,536 tokens and refused it (its 413, code 5021). The message carries Cloudflare's words, for example upstream workers-ai returned 413: The estimated number of input and maximum output tokens (130569) exceeded this model context window limit (65536). metadata.retryable is false: make the state shorter. It costs nothing.
503model_unavailableSage: Levanto refused our account (its 401, 402 or 403, for example when our month's usage is used up); the message says the model is out of capacity at the provider. GLiNER: Fastino refused our account for credit or billing (its 402 or 403). The message says the model is out of capacity at the provider. Nimble: Bespoke refused our account for credit (its 402); the message says the model is out of capacity at the provider, and metadata.retryable is false. Clef and Clef-flash: Cloudflare refused our account, for our token or plan (its 401 or 403), or because the account's free daily allocation is used up (its 429 with code 3036, until 00:00 UTC). The message says the model is out of capacity at the provider, and metadata.retryable is false. A 401, 402 or 403 from any System One provider (TypeSafe's or OpenRouter's, on Jev or Mercury Decide, and RouterPlus's, on Decider 2B and Kev 4B, and Perplexity's, on Perplexity Decider, too) reads the same: it is our account, never yours. On a connection with your own provider key, that provider's 401, 402 or 403 is your account, and it keeps its status. Other models are not affected.
503gateway_error"all deployments cooling down", with retry-after: 5: the deployment's health circuit is open. On Sage, GLiNER, Nimble, Clef, Clef-flash, the two RouterPlus models and Perplexity Decider this is also what most requests see while the provider refuses our account: each refusal counts against the provider's circuit, so after two of them only one request in each 30 s reaches the provider and gets the message above, and the rest get this one. Decider 2B and Kev 4B share one deployment and so one circuit: two RouterPlus 5xx in a row open it for both. A provider's 429 never counts, so Mercury Decide's daily cap on OpenRouter, Nimble's limit of 8 requests at once, Cloudflare's busy answer on Clef, RouterPlus's limit of about 200 a second and Perplexity's limit of 10 a second stay 429s (below), and Cloudflare's code 3036 stays the 503 above.
422 and other 4xxupstream_errorThe provider refused the request as wrong. The message carries the provider's own reason, and x-tm-upstream-status its status. The gateway does not send a request the provider called wrong to another provider. A Fastino 404 (an unknown upstream model id: our misconfiguration, never yours) is a 502 upstream_error, with x-tm-upstream-status: 404. A Nimble request above Nimble's own limits (more than 64 questions, a choice with one option, or a question's prompt above 32,768 tokens with the state) is Bespoke's 422, relayed here with Bespoke's words; it costs nothing. A Clef or Clef-flash request outside Clef's own limits (more than 64 questions, a question id that does not match ^[A-Za-z0-9_.-]{1,100}$, or a choice with one option) is Cloudflare's 422 (code 5012) or 400 (code 5006), relayed here with Cloudflare's words, for example upstream workers-ai returned 422: Request body failed validation: questions: Dictionary should have at most 64 items after validation, not 65; it costs nothing. A request that RouterPlus refuses (Decider 2B, Kev 4B) is RouterPlus's 400, relayed here with RouterPlus's reason and x-tm-upstream-status: 400; it costs nothing. RouterPlus refuses a malformed question (invalid questions, then the reason for each question id, for example score needs criteria: a list of 2 to 10 levels) and a Kev 4B request above 8,192 tokens (the question's id, then ContextOverflow: branch too long). A Perplexity Decider request with more than 128 questions is Perplexity's 400, relayed here with Perplexity's words (Each request needs between 1 and 128 questions); it costs nothing.
429rate_limitMercury Decide: more than its pool allows (20 requests a minute in all, 15 per organization; 5 in flight at once, 3 per organization; 700,000 tokens a minute, 525,000 per organization), with retry-after — the gateway refuses before OpenRouter does. x-tm-limit-kind names the limit (rpm, tpm, concurrency, or cooldown: the pool's pause after a relayed OpenRouter 429, 1 to 60 s) and x-tm-limit-id says whether it was the pool (pool:…) or your organization's share (pool-share:…). Or OpenRouter's daily cap on free requests is used up (1,000 a day across our account): OpenRouter's 429, relayed with its retry-after and x-tm-upstream-status: 429, until its daily reset; it never opens the deployment's circuit, and it pauses the pool for at most 60 s.
429rate_limitNimble: more than its pool allows (7 in flight at once, 5 per organization; 600 requests a minute, 450 per organization; 15,000,000 tokens a minute, 11,250,000 per organization), with retry-after, before anything reaches Bespoke; x-tm-limit-kind names the limit (rpm, tpm, concurrency, or cooldown). Or Bespoke already runs 8 requests for our account: Bespoke's 429, relayed with its Retry-After and x-tm-upstream-status: 429; it never opens the deployment's circuit, and it pauses the pool for at most 60 s.
429rate_limitClef and Clef-flash: more than their shared pool allows (200 requests a minute, 150 per organization; at most 8 requests at a time, 6 per organization; 2,000,000 tokens a minute, 1,500,000 per organization), with retry-after, before anything reaches Cloudflare; x-tm-limit-kind names the limit (rpm, tpm, concurrency, or cooldown). Or Workers AI is busy: Cloudflare's 429 with code 3040, relayed with Cloudflare's words (upstream workers-ai returned 429: Capacity temporarily exceeded, please try again.), its Retry-After when it sends one, and x-tm-upstream-status: 429. It never opens the deployment's circuit, and it pauses the pool for at most 60 s (1 s when Cloudflare sends no Retry-After). Any other Cloudflare 429 is relayed the same way, except code 3036 (the 503 above).
429rate_limitDecider 2B and Kev 4B: more than their shared pool allows (11,400 requests a minute, 8,550 per organization; 15,000,000 tokens a minute, 11,250,000 per organization; 59 in flight at once, 44 per organization), or more than your organization's own limits, which bind first (by default 300 requests a minute and 8 in flight), with retry-after, before anything reaches RouterPlus; x-tm-limit-kind names the limit (rpm, tpm, concurrency, or cooldown). Or RouterPlus is at its limit of about 200 requests a second for the whole endpoint: RouterPlus's 429, relayed with its retry-after and x-tm-upstream-status: 429; it never opens the deployment's circuit, and it pauses the pool for at most 60 s, for both models.
429rate_limitPerplexity Decider: more than its pool allows (540 requests a minute, 405 per organization; at most 8 requests at a time, 6 per organization; 15,000,000 tokens a minute, 11,250,000 per organization), or more than your organization's own limits, which bind first (by default 300 requests a minute, 1,000,000 tokens a minute and 8 in flight), with retry-after, before anything reaches Perplexity; x-tm-limit-kind names the limit (rpm, tpm, concurrency, or cooldown). The state is held once per question, so a long state with many questions can be larger than a token limit on its own (x-tm-limit-kind: tpm): ask fewer questions per call, or send a shorter state. Or Perplexity is at its limit of 10 requests a second for our organization: Perplexity's 429, relayed with its words (Request rate limit exceeded, please try again later.), its Retry-After and x-tm-upstream-status: 429; it never opens the deployment's circuit, and it pauses the pool for at most 60 s.
504upstream_errorA System One provider answered 408, its own timeout. On Clef and Clef-flash the message is upstream workers-ai returned 408: Request timeout, with x-tm-upstream-status: 408. Retry after a few seconds. A 408 is not tried on a second provider, and it does not count against the deployment's circuit. The gateway's own wait for a decision (60 s) also ends in a 504 (next row).
429, 5xx (and 529)rate_limit, upstream_errorEvery provider failed. A rate limit, an outage or an overload at one provider is tried on the next one first. Levanto answers 503 while Sage loads: retry after a few seconds. Fastino answers 425 or 503 while GLiNER starts (a cold start), and a cold start can also outlast the gateway's wait for a decision (60 s, a 504): both are a retryable upstream_error, so retry after a few seconds. A Fastino 429 is a rate_limit. Bespoke answers 503 with Retry-After while Nimble starts, 529 when it is busy and 502 when the model fails: each is a retryable upstream_error that counts against the deployment's circuit (two in a row open it for 30 s), so retry after a second or two. Cloudflare answers 500 when Clef fails ("Model execution failed"): a retryable upstream_error that counts against the circuit in the same way. A RouterPlus 5xx is a retryable upstream_error too, and it counts against the one circuit of Decider 2B and Kev 4B: two in a row open it for 30 s for both. A RouterPlus cold start (15 to 20 s) is not an error: it fits inside the gateway's 60 s wait and answers 200. Perplexity answers 504 when its model does not answer in about a minute (an HTML page, so the message has no words of Perplexity's): a retryable upstream_error that counts against the circuit in the same way.
401, 402, 403authEvery provider refused our own account with it. This is never your key's fault; it is tried on the next provider first. For GLiNER a 402 or 403, and for every System One model (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage) a 401, 402 or 403, is the 503 above.

Jev reads at most 64 000 tokens per request through TypeSafe (32 000 for the state plus the longest question), and 32 000 through OpenRouter. Mercury Decide reads at most 32 768 tokens per request. Sage reads at most 32 768 tokens per request, the state and every question together. Nimble reads at most 32,768 tokens for each question's prompt, the state included, and never cuts a prompt. Clef and Clef-flash read at most 65,536 tokens per request, the state and every question together, as Cloudflare estimates them (about 4 characters a token); Cloudflare refuses a longer request with the 413 context_overflow above. Decider 2B's context is 25,600 tokens and Kev 4B's 8,192. GLiNER-2.5-Decide's context is 8,192 tokens. A request above a provider's limit is refused by that provider. Decider 2B cuts it instead, and bills only the tokens the model read; Kev 4B refuses it with a 400. See Decider 2B and Kev 4B. Perplexity Decider reads at most 262,144 tokens for each question's prompt (the state and that question), and never cuts one: Perplexity refuses a longer one with the 400 context_overflow above. See Perplexity Decider v1 27B.

Examples

Jev

curl
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"typesafe/jev-1.13",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
python
No SDK has a decisions method, so send plain HTTP
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "typesafe/jev-1.13",
        "state": "Help! My payouts have been failing for 3 days.",
        "questions": {"team": {"type": "choice", "instructions": "Which team?",
                               "criteria": {"billing": "Payments", "technical": "Bugs"}}},
    },
)
answer = r.json()["answers"]["team"]
print(answer["choice"], answer["confidence"])
javascript
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "typesafe/jev-1.13",
    state: { ticket: "Help! My payouts have been failing for 3 days." },
    questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } },
  }),
});
const { answers, usage } = await r.json();
console.log(answers.urgent.noul, usage.cost);

Mercury Decide

The same System One body as for Jev, under Mercury Decide's id. A 429 rate_limit here is either our pool (wait retry-after) or OpenRouter's daily cap on free requests (wait for its reset): switch on the class, as Errors says, and read the message.

curl
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"inception/mercury-decide",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
python
Plain HTTP; the answers read exactly as Jev's
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "inception/mercury-decide",
        "state": "Help! My payouts have been failing for 3 days.",
        "questions": {
            "team": {"type": "choice", "instructions": "Which team?",
                     "criteria": {"billing": "Payments", "technical": "Bugs"}},
            "angry": {"type": "score", "instructions": "How angry is the customer?",
                      "criteria": ["Calm", "Mild", "Moderate", "Upset", "Furious"]},
        },
    },
)
if r.status_code == 429:
    print("rate limited:", r.json()["error"]["message"], "retry after", r.headers.get("retry-after"))
else:
    answers = r.json()["answers"]
    print(answers["team"]["choice"], answers["angry"]["score"], r.json()["usage"]["cost"])  # cost is 0
javascript
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "inception/mercury-decide",
    state: { ticket: "Help! My payouts have been failing for 3 days." },
    questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } },
  }),
});
const { answers, usage } = await r.json();
console.log(answers.urgent.noul, usage.cost); // 0.95, 0

Nimble

The same System One body as for Jev, under Nimble's id. Keep to Nimble's limits: 1 to 64 questions, and 2 or more options in a choice.

curl
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"bespokelabs/nimble-v3",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
python
Plain HTTP; the answers read exactly as Jev's, at full precision
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "bespokelabs/nimble-v3",
        "state": "Help! My payouts have been failing for 3 days.",
        "questions": {
            "team": {"type": "choice", "instructions": "Which team?",
                     "criteria": {"billing": "Payments", "technical": "Bugs"}},
            "urgent": {"type": "noul", "instructions": "Is this urgent?"},
        },
    },
)
answers = r.json()["answers"]
print(answers["team"]["choice"], answers["urgent"]["noul"])  # e.g. billing 0.997817283712868
javascript
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "bespokelabs/nimble-v3",
    state: { ticket: "Help! My payouts have been failing for 3 days." },
    questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } },
  }),
});
const { answers, usage } = await r.json();
console.log(answers.urgent.noul, usage.input_tokens, usage.cost);

Clef

The same System One body as for Jev, under Clef's or Clef-flash's id. Keep to Clef's limits: 1 to 64 questions, ids of 1 to 100 letters, digits, _, . or -, and 2 or more options in a choice.

curl
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"cloudflare/clef",
       "state":"Checkout has been failing for every customer for the last hour.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this support request urgent?"}}}'
python
Plain HTTP; the answers read as Jev's, to 4 decimals
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "cloudflare/clef",
        "state": "Checkout has been failing for every customer for the last hour.",
        "questions": {
            "urgent": {"type": "noul", "instructions": "Is this support request urgent?"},
            "team": {"type": "choice", "instructions": "Which team should handle this request?",
                     "criteria": {"billing": "Payments and invoices", "technical": "Bugs and outages",
                                  "sales": "Plans and upgrades"}},
            "severity": {"type": "score", "instructions": "How severe is the customer impact?",
                         "criteria": ["No impact", "Minor", "Major", "Critical"]},
        },
    },
)
answers = r.json()["answers"]
team = answers["team"]
print(team["choice"], team["probabilities"][team["choice"]], team["confidence"])  # e.g. technical 0.8274 0.5654
javascript
Clef-flash, the smaller model
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "cloudflare/clef-flash",
    state: { ticket: "Checkout has been failing for every customer for the last hour." },
    questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } },
  }),
});
const { answers, usage } = await r.json();
console.log(answers.urgent.noul, usage.input_tokens, usage.output_tokens); // output_tokens is always 0

RouterPlus models

The same System One body as for Jev, under Decider 2B's id. For Kev 4B, change only model. A first request after a quiet period can take 15 to 20 s while the model starts, so give your client a timeout above 60 s, the gateway's own wait.

curl
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"routerplus/decider-2b",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
python
Plain HTTP, with a timeout that covers a cold start
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "routerplus/decider-2b",
        "state": "Help! My payouts have been failing for 3 days.",
        "questions": {
            "team": {"type": "choice", "instructions": "Which team?",
                     "criteria": {"billing": "Payments", "technical": "Bugs"}},
            "urgent": {"type": "noul", "instructions": "Is this urgent?"},
        },
    },
    timeout=70,
)
answers = r.json()["answers"]
print(answers["team"]["choice"], answers["urgent"]["noul"], r.headers["x-tm-upstream-model"])  # e.g. billing 0.97 decider-2b
javascript
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "routerplus/decider-2b",
    state: { ticket: "Help! My payouts have been failing for 3 days." },
    questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } },
  }),
  signal: AbortSignal.timeout(70_000),
});
const { answers, usage } = await r.json();
console.log(answers.urgent.noul, usage.input_tokens, usage.cost);

Perplexity

The same System One body as for Jev, under Perplexity Decider's id. Keep to its limits: 1 to 128 questions, and text or JSON only. Each question is billed with the whole state, so ask only the questions you need.

curl
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"perplexity/pplx-decider-v1-27b",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
python
Plain HTTP; the answers read as Jev's, at full precision
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "perplexity/pplx-decider-v1-27b",
        "state": "Help! My payouts have been failing for 3 days.",
        "questions": {"urgent": {"type": "noul", "instructions": "Is this urgent?"}},
    },
)
body = r.json()
print(body["answers"]["urgent"]["noul"], body["usage"]["input_tokens"])  # e.g. 0.9297849172276498 95
javascript
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "perplexity/pplx-decider-v1-27b",
    state: { ticket: "Help! My payouts have been failing for 3 days." },
    questions: {
      team: { type: "choice", instructions: "Which team should handle `ticket`?", criteria: { billing: "Payments", technical: "Bugs" } },
      urgent: { type: "noul", instructions: "Is `ticket` urgent?" },
    },
  }),
});
const { answers, usage } = await r.json();
// Two questions: Perplexity bills the state twice.
console.log(answers.team.choice, answers.urgent.noul, usage.input_tokens, usage.cost);

Sage

The same System One body as for Jev, under Sage's id. Each question is billed with the whole state, so ask only the questions you need.

bash
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"levanto/sage-1.2",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'

GLiNER

curl
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"fastino/gliner-2.5-decide",
       "state":"Help! My payouts have been failing for 3 days.",
       "schema":{"classifications":[{"task":"Is this urgent?","labels":["yes","no"]}]}}'
python
Plain HTTP; each answer sits under its task name
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "fastino/gliner-2.5-decide",
        "state": "Help! My payouts have been failing for 3 days.",
        "schema": {"classifications": [{"task": "team",
                                        "labels": ["billing: payments", "technical: bugs"]}]},
    },
)
answer = r.json()["answers"]["team"]
print(answer["label"], answer["confidence"])  # a confidence, not a calibrated probability
javascript
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "fastino/gliner-2.5-decide",
    state: "Help! My payouts have been failing for 3 days.",
    schema: {
      classifications: [
        { task: "topics", labels: ["billing", "outage", "fraud"], multi_label: true, cls_threshold: 0 },
      ],
    },
  }),
});
const { answers, usage } = await r.json();
console.log(answers.topics, usage.input_tokens, usage.output_tokens, usage.cost);

Model discovery

bash
curl -s "https://api.routerplus.com/v1/models?output_modalities=decisions" \
  -H "Authorization: Bearer $TM_API_KEY"

Decision models carry architecture.output_modalities: ["decisions"] in GET /v1/models, and output_modalities: ["decisions"] in the public feed https://app.routerplus.com/api/models.json. The Anthropic shape of GET /v1/models never lists them.

Markdown source for agents: /docs/api-decisions.md · index at /llms.txt