# POST /v1/decisions

Ask a decision model typed questions about a state, and get one typed answer per question.
A decision model does not write text. You send a **state** (a message, a ticket, a record)
and your **questions**. Each answer is a choice, a score, a yes/no, an order, a set of tags
or extracted fields, with the probability or confidence behind it. Use it for routing,
classification, moderation and other decision points in your code.

The catalog has ten decision models:

| Model | Catalog id | Made by | Providers | Question format |
|---|---|---|---|---|
| Jev 1.13 | `typesafe/jev-1.13` | TypeSafe | `typesafe`, then `openrouter-decisions` | System One `questions`: `type` is `choice`, `score` or `noul` |
| Mercury Decide | `inception/mercury-decide` | Inception | `openrouter-decisions-free` | System One `questions`, exactly as for Jev |
| Bespoke Nimble v3 | `bespokelabs/nimble-v3` | Bespoke Labs | `bespokelabs` | System One `questions`, as for Jev, within Nimble's own limits |
| Clef | `cloudflare/clef` | Cloudflare | `workers-ai` | System One `questions`, as for Jev, within Clef's own limits |
| Clef-flash | `cloudflare/clef-flash` | Cloudflare | `workers-ai` | System One `questions`, as for Jev, within Clef's own limits |
| Decider 2B | `routerplus/decider-2b` | RouterPlus | `routerplus` | System One `questions`, exactly as for Jev |
| Kev 4B | `routerplus/kev-4b` | RouterPlus | `routerplus` | System One `questions`, exactly as for Jev |
| Perplexity Decider v1 27B | `perplexity/pplx-decider-v1-27b` | Perplexity | `perplexity-decisions` | System One `questions`, as for Jev, within its own limits; it bills the state once per question |
| Sage 1.2 | `levanto/sage-1.2` | Levanto | `levanto` | System One `questions`, as for Jev, within Sage's own limits; it bills the state once per question |
| GLiNER-2.5-Decide | `fastino/gliner-2.5-decide` | Fastino | `fastino` | A GLiNER `schema`: `classifications`, `entities`, `structures`, `relations` |

**The question format follows the model.** Every call has the same envelope: `model` and
`state` in, `answers` and `usage` out. The questions go in the model's own form: a
`questions` map for the System One models, a `schema` object for GLiNER. What
goes inside a question, and what comes back inside an answer, is the model's own format.
Nine models share one format, **System One** (TypeSafe's): Jev, Mercury Decide, Nimble,
Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage. A body written for one runs on
the others with only `model` changed, when it is inside each model's limits. A question in
another model's format is a 400. The ten are different models, so the gateway never fails
over from one to another.

The gateway calls these endpoints:

| Model | Order | Provider | Endpoint the gateway calls | Model id sent upstream |
|---|---|---|---|---|
| Jev | 1 | TypeSafe (direct API) | `POST https://api.typesafe.ai/v1/systemone` | `jev-1.13.0` |
| Jev | 2 | OpenRouter | `POST https://openrouter.ai/api/alpha/decisions` | `typesafe/jev-1.13-20260917` |
| Mercury Decide | 1 | OpenRouter | `POST https://openrouter.ai/api/alpha/decisions` | `inception/mercury-decide:free` |
| Nimble | 1 | Bespoke Labs (direct API) | `POST https://api.bespokelabs.ai/v1/nimble/systemone` | `nimble-v3` |
| Clef | 1 | Cloudflare Workers AI (direct API) | `POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef` | `clef` |
| Clef-flash | 1 | Cloudflare Workers AI (direct API) | `POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef-flash` | `clef-flash` |
| Decider 2B | 1 | RouterPlus (our own service, direct API) | `POST /v1/decisions` on RouterPlus's endpoint | `decider-2b` |
| Kev 4B | 1 | RouterPlus (our own service, direct API) | `POST /v1/decisions` on RouterPlus's endpoint | `kev-4b` |
| Perplexity Decider | 1 | Perplexity (direct API) | `POST https://api.perplexity.ai/v1/decisions` | `pplx-decider-v1-27b` |
| Sage | 1 | Levanto (direct API) | `POST https://sage.levanto.ai/v1/systemone` | `sage-latest` |
| GLiNER | 1 | Fastino (direct API) | `POST https://api.fastino.ai/v1/chat/completions` | `fastino/GLiNER-2.5-Decide` |

For Jev you send one request shape. The gateway sends each provider the id it expects, and
you get one answer shape back, whichever provider served it. If TypeSafe fails before it
answers (overloaded, down, rate limited), the gateway tries OpenRouter. Mercury Decide,
Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sage and GLiNER have one
provider each, so their requests have no second provider to try. On Workers AI, `{account_id}` is our
Cloudflare account, and the path names the model. Decider 2B and Kev 4B share one
RouterPlus deployment: one pool and one health circuit serve both. The `x-tm-provider`
header names the provider that answered.

To compare the models without code, use **Decision** in the [playground](/docs/playground). It
takes one set of questions and translates it for each model.

## Authentication

| Header | Format |
|---|---|
| `Authorization` | `Bearer tm_vk_...` |
| `x-api-key` | `tm_vk_...` — also accepted, same key |

A missing or invalid key returns 401 with `error_type: "auth"`.

## Request

JSON body, at most 1 MB (413 `request_too_large` above that). The gateway checks every rule
on this page **before** it reserves money or calls a provider. A refused request has no
ledger row, no `x-tm-attempts` header and no charge.

| Parameter | Type | Behavior |
|---|---|---|
| `model` | string, required | A decision model id. A dated id of a listed decision model is accepted and routes as the model: `typesafe/jev-1.13-20260917` (OpenRouter's dated id) as `typesafe/jev-1.13`, `inception/mercury-decide-20260930` as `inception/mercury-decide`. OpenRouter's variant suffix (`inception/mercury-decide:free`) is not an id here: send `inception/mercury-decide`. Bespoke's own names (`nimble-v3`, `nimble-latest`) are not ids here either: send `bespokelabs/nimble-v3`. Nor are Cloudflare's (`clef`, `clef-flash`, `@cf/cloudflare/clef`, `@cf/cloudflare/clef-flash`): send `cloudflare/clef` or `cloudflare/clef-flash`. RouterPlus's own short names (`decider-2b`, `kev-4b`) are not ids here: send `routerplus/decider-2b` or `routerplus/kev-4b`. Perplexity's own name (`pplx-decider-v1-27b`) is not an id here either: send `perplexity/pplx-decider-v1-27b`. A chat, image or video model is a 400 that names its route. An unknown id is a 404. |
| `state` | required | What the model evaluates. Its forms are the model's: see [System One questions](#system-one-questions) and [GLiNER schema](#gliner-schema) below. |
| `questions` | object, the System One models, required | A map of question id to question. You choose the ids; each answer comes back under the same id. At least one question. On GLiNER, `questions` is a 400 that says the model takes a GLiNER `schema`. |
| `schema` | object, GLiNER only, required | GLiNER's own schema. See [GLiNER schema](#gliner-schema). On Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider or Sage, `schema` is a 400 that says the model takes `questions`. |
| `threshold`, `include_confidence`, `include_spans` | GLiNER only | See [GLiNER options](#gliner-options). |
| `provider` | object | Routing controls, read by the gateway and never forwarded: `only`, `order`, `ignore`, `require_parameters` and the rest — see [Routing policies](/docs/routing-policies). The provider ids are `typesafe` and `openrouter-decisions` for Jev, `openrouter-decisions-free` for Mercury Decide, `bespokelabs` for Nimble, `workers-ai` for Clef and Clef-flash (not `cloudflare`, which is an OpenRouter host tag), `routerplus` for Decider 2B and Kev 4B, `perplexity-decisions` for Perplexity Decider (not `perplexity`, which is an OpenRouter host tag), `levanto` for Sage and `fastino` for GLiNER. |

### System One questions

Eight models take the same question format, TypeSafe's **System One**: Jev, Mercury Decide,
Nimble, Clef, Clef-flash, Decider 2B, Kev 4B and Perplexity Decider. Everything in this
section holds for each of the eight; only `model` differs, and Nimble, Clef, Clef-flash, the
two RouterPlus models and Perplexity Decider have some limits of their own (see
[Bespoke Nimble v3](#bespoke-nimble-v3), [Clef and Clef-flash](#clef-and-clef-flash),
[Decider 2B and Kev 4B](#decider-2b-and-kev-4b) and
[Perplexity Decider v1 27B](#perplexity-decider-v1-27b)). The examples show Jev; put
another System One model's id in `model`, for example `inception/mercury-decide`,
`bespokelabs/nimble-v3`, `cloudflare/clef`, `routerplus/decider-2b` or
`perplexity/pplx-decider-v1-27b`, and the same body runs on that model.

**State.** A non-empty string, or any JSON object or array (a chat log, a record,
application state). Questions can point at a field by name: `` "Is `ticket.body` urgent?" ``.

**Questions.** Every question has a `type`, `instructions` and, for most types, `criteria`.
`instructions` is a non-empty string, or a JSON object or array: put the question in one
field and the data it refers to in others.

| `type` | `criteria` | The answer |
|---|---|---|
| `choice` | Required. An object of option name to description, 1 to 255 options. A description can be a string, JSON, or `null` when the name says enough. | `choice` (the most likely option), `probabilities` (every option), `confidence` |
| `score` | Required. An ordered array of level descriptions, low to high, 2 to 10 levels. | `score` (probability-weighted, can fall between levels), `legend`, `probabilities`, `confidence` |
| `noul` | Optional. `{"true": "...", "false": "..."}` — what yes and no mean. | `noul`: the probability that the answer is yes, 0 to 1 |

A question takes only `type`, `instructions` and `criteria`. Any other field in a question
is a 400 that names it. A question with `kind` (the format of Levanto's own `/decide` API)
is a 400 that says the model takes System One questions. Ask many questions in one call: the
model reads the state once and answers every question against it, so one call with ten
questions is cheaper and faster than ten calls. Perplexity Decider and Sage are the
exceptions: they read and bill the state once per question, so on them one call saves round
trips, not tokens (see [Perplexity Decider v1 27B](#perplexity-decider-v1-27b) and
[Sage 1.2](#sage-12)).

```json
{
  "model": "typesafe/jev-1.13",
  "state": { "ticket": "Help! My payouts have been failing for 3 days." },
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle `ticket`?",
      "criteria": { "billing": "Payments, invoices, refunds", "technical": "Bugs, outages, integrations", "sales": null }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated is the customer?",
      "criteria": ["Calm", "Annoyed", "Angry"]
    },
    "urgent": {
      "type": "noul",
      "instructions": "Is this urgent?",
      "criteria": { "true": "Explicitly time-sensitive", "false": "No urgency expressed" }
    }
  }
}
```

#### Mercury Decide

Inception's Mercury Decide is a System One model: the state, the questions and the answers
above are its format too. What is its own:

| | Mercury Decide |
|---|---|
| Catalog id | `inception/mercury-decide`. The dated `inception/mercury-decide-20260930` routes as the same model. |
| Route | OpenRouter's decisions endpoint only (`POST https://openrouter.ai/api/alpha/decisions`, provider id `openrouter-decisions-free`), as `inception/mercury-decide:free`. Inception's own API does not serve it yet. |
| Price | $0 in and $0 out: OpenRouter serves only the model's free variant today, and the gateway passes that through. `usage.cost` is `0`. See [Billing](#billing). |
| Context | 32,768 tokens, the state and every question together. |
| Our limit | The deployment's pool takes 20 requests a minute in all, and each organization at most 15 of them; at most 5 requests in flight at once, 3 per organization; and 700,000 tokens a minute in all, 525,000 per organization. Tokens are held at the gateway's estimate of the body while a request runs (a state near the 32k context is held at about 44,000) and corrected to OpenRouter's count when it settles (about 33,000 for that state), so the 15-per-organization share of requests is the limit that binds, not tokens. A burst above any of these is a 429 `rate_limit` from the gateway, with `retry-after` and the `x-tm-limit-*` headers, before anything reaches OpenRouter: `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown` for the pause after a relayed OpenRouter 429, next row) and `x-tm-limit-id` says whether it was the pool (`pool:…`) or your organization's share of it (`pool-share:…`). |
| OpenRouter's limit | A free variant shares OpenRouter's daily cap on free requests across our whole account (1,000 a day). When that is used up, OpenRouter answers 429 until its daily reset, and the gateway relays it as 429 `rate_limit` with OpenRouter's `retry-after` and `x-tm-upstream-status: 429`. A relayed 429 pauses the deployment's pool for that `retry-after`, at least 1 s and at most 60 s: a request inside the pause gets the gateway's own 429 `rate_limit`, with `x-tm-limit-kind: cooldown`, `x-tm-limit-id: pool:…` and no `x-tm-upstream-status`. After the pause, requests reach OpenRouter again and meet the cap again, until its daily reset. A relayed 429 never counts against the deployment's health circuit, so the daily cap is 429s until the reset, not the 503 of a cooling-down deployment. The gateway does not count the daily cap itself, so a request can pass our pool and still meet it. |
| Retention | OpenRouter publishes no retention terms for this free endpoint, so the deployment declares `prompt_logging: "retained"`. See [Data policy](/docs/data-policy). |

A free endpoint is, in OpenRouter's own words, not production-suitable. For a decision that
must come back, keep Jev as your System One fallback in your own code: the gateway never
fails over between models.

#### Bespoke Nimble v3

Bespoke Labs' Nimble is a System One model: the state, the questions and the answers above
are its format too. The gateway calls Bespoke's own API. What is its own:

| | Bespoke Nimble v3 |
|---|---|
| Catalog id | `bespokelabs/nimble-v3`. Bespoke's names `nimble-v3` and `nimble-latest` are not ids here. |
| Route | Bespoke Labs' API only (`POST https://api.bespokelabs.ai/v1/nimble/systemone`, provider id `bespokelabs`), as `nimble-v3`. |
| Price | $0.04 per million input tokens. Output is free. This is Bespoke's own price, with no markup. See [Billing](#billing). |
| Questions | 1 to 64 questions per call. A `choice` takes 2 to 255 options: a choice with one option, which Jev takes, is refused. A `score` takes 2 to 10 levels through the gateway, as for every System One model (Nimble itself takes up to 255). The gateway does not check Nimble's own limits: a body above them is Bespoke's 422, relayed as an `upstream_error`, and costs nothing. (The Playground's Decision mode does check them, before a run.) |
| Structured content | `instructions` and a `choice` option's description may be a JSON object or array, as on Jev. Bespoke's published schema types them as strings, but Nimble takes them: a probe with both answered 200 (2026-10-02). |
| Context | 32,768 tokens for each question's prompt, the state included. Nimble never cuts a prompt: a longer one is refused. Bespoke counts the state and the questions once per call, not once per question. |
| Busy | Bespoke answers 503 with `Retry-After` while Nimble starts, 529 when it is busy (Bespoke says: retry after about one second) and 502 when the model fails. Each is a retryable `upstream_error` that counts against the deployment's health circuit, like every 5xx: two in a row open it for 30 s, and in that window every request but one probe gets 503 `gateway_error`, "all deployments cooling down". Retry after a second or two, with backoff. The gateway does not relay Bespoke's `Retry-After` on a 503. See [Errors](#errors). |
| Numbers | Full precision: Jev's numbers have 2 decimals, Nimble's are not rounded. A very small probability can come back in exponent form, for example `3.7e-06`. `confidence` is 1 when one option has all the probability and 0 when every option is equally likely. It is not the chance that the answer is right. |
| Our limit | The deployment's pool takes 7 requests in flight at once, and each organization at most 5 of them (Bespoke allows our account 8; one is kept for our test environment). It also takes 600 requests a minute (450 per organization) and 15,000,000 tokens a minute (11,250,000 per organization). A burst above any of these is a 429 `rate_limit` from the gateway, with `retry-after` and the `x-tm-limit-*` headers, before anything reaches Bespoke; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown` for the pause after a relayed Bespoke 429, next row). |
| Bespoke's limit | Bespoke runs at most 8 requests at once for our whole account. Above that it answers 429 with `Retry-After`. The gateway relays it as 429 `rate_limit` with `x-tm-upstream-status: 429`, and pauses the deployment's pool for that `retry-after`, at least 1 s and at most 60 s. A relayed 429 never counts against the deployment's health circuit. |
| Retention | Bespoke keeps request content for up to 30 days, to debug the service and check for abuse, and then deletes it. The deployment declares `prompt_logging: "retained"`. Bespoke does not train on API content unless an organization opts in; ours does not. See [Data policy](/docs/data-policy). |

#### Clef and Clef-flash

Cloudflare's Clef and Clef-flash are System One models: the state, the questions and the
answers above are their format too. Clef-flash is the smaller model (9 billion parameters,
against Clef's 27 billion): it is faster and costs less. The gateway calls Cloudflare's own
API, Workers AI. What is their own:

| | Clef and Clef-flash |
|---|---|
| Catalog ids | `cloudflare/clef` and `cloudflare/clef-flash`. Cloudflare's names `clef`, `clef-flash`, `@cf/cloudflare/clef` and `@cf/cloudflare/clef-flash` are not ids here. |
| Route | Cloudflare Workers AI's REST API only, provider id `workers-ai`: `POST https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef` as `clef`, and `.../@cf/cloudflare/clef-flash` as `clef-flash`. `{account_id}` is our Cloudflare account. The path picks the model, and the body's `model` must be the same model; the gateway sets both. Cloudflare puts each answer inside its own envelope (`result`, `success`, `errors`, `messages`). The gateway opens it, so you get Jev's answer shape. A 200 whose envelope does not say `success: true` is a 502 `upstream_error` and costs nothing. |
| Price | Clef: $0.24 per million input tokens. Clef-flash: $0.09 per million input tokens. Output is free: Cloudflare reports `output_tokens: 0` on every answer. This is Cloudflare's own price, with no markup. See [Billing](#billing). |
| Questions | 1 to 64 questions per call. Each question id (a key of `questions`) is 1 to 100 letters, digits, `_`, `.` or `-`: it must match `^[A-Za-z0-9_.-]{1,100}$`. A `choice` takes 2 to 255 options: a choice with one option, which Jev takes, is refused. A `score` takes 2 to 10 levels. The gateway does not check Clef's own limits: a body outside them is Cloudflare's 422 (or 400), relayed as an `upstream_error` with Cloudflare's words, and costs nothing. (The Playground's Decision mode does check them, before a run.) |
| Structured content | `instructions`, a `score` level and a `choice` option's description may be a JSON object or array, as on Jev, and an option's description may be `null`. Cloudflare's schema says so. A probe with a JSON object as `instructions` and JSON option descriptions answered 200 (2026-10-02). |
| Context | 65,536 tokens, the state and every question together. Before the model runs, Cloudflare estimates the request's tokens at about 4 characters a token. It refuses a request estimated above 65,536 tokens (about 262,000 characters) with a 413, which the gateway returns as 413 `context_overflow`; it costs nothing. Cloudflare's schema says that a long state is cut to fit, but in our test (2026-10-02) a request of 522,286 characters was refused, not cut. |
| Images | Not accepted through the gateway. Clef's `images` field, Cloudflare's addition to System One, is dropped and recorded in `x-tm-dropped-params`, never forwarded. With `"provider": {"require_parameters": true}` it is a 400. |
| `request_id` | Do not send a top-level `request_id`. The gateway drops and records it, and never forwards it: Workers AI reads that field as a lookup in its queue of async requests, and answers 404. |
| Numbers | 4 decimals: Jev's numbers have 2, and Nimble's are not rounded. `confidence` is on Clef's own scale, not on Jev's. Cloudflare says only that it is derived from the probabilities. On one live answer (a choice of three options, the top one at 0.8274), Clef's `confidence` was 0.5654, where Jev's formula gives 0.7411. So a threshold that you tuned on Jev's `confidence` does not transfer to Clef: tune it on Clef, or use the probabilities. `confidence` is not the chance that the answer is right. |
| Busy | When Workers AI is busy, Cloudflare answers 429 with its code 3040, "Capacity temporarily exceeded, please try again." The gateway relays it as 429 `rate_limit`, with Cloudflare's words and `x-tm-upstream-status: 429`, and the deployment's pool pauses for Cloudflare's `Retry-After`, at least 1 s and at most 60 s. Without a `Retry-After`, the relayed 429 has no `retry-after` either, and the pool pauses for 1 s. A relayed 429 never counts against the deployment's health circuit. Retry after a second or two, with backoff. A 5xx (for example Cloudflare's 500, "Model execution failed") is a retryable `upstream_error` that counts against the circuit, like every 5xx: two in a row open it for 30 s. A 408, Cloudflare's own timeout, is a retryable 504 `upstream_error` that does not count. See [Errors](#errors). |
| Our limit | Clef and Clef-flash share one pool. It takes 200 requests a minute (150 per organization), at most 8 requests at a time (6 per organization) and 2,000,000 tokens a minute (1,500,000 per organization). A burst above any of these is a 429 `rate_limit` from the gateway, with `retry-after` and the `x-tm-limit-*` headers, before anything reaches Cloudflare; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown` for the pause after a relayed Cloudflare 429). |
| Cloudflare's limit | Cloudflare publishes 300 requests a minute for Clef's task type, Text Generation. It does not say whether Clef and Clef-flash share that limit, so our pool stays below it. Cloudflare's 429 with code 3036 says that our account's free daily allocation is used up. The gateway answers it with 503 `model_unavailable` and `metadata.retryable: false`: that is our account, never yours. It never counts against the deployment's health circuit. |
| Retention | Cloudflare says that it does not read, store or train on requests to Clef, or on their responses. The deployment declares `prompt_logging: "none"`. The gateway calls Workers AI directly, never through Cloudflare's AI Gateway, which logs prompts and responses by default. See [Data policy](/docs/data-policy). |

#### Decider 2B and Kev 4B

RouterPlus is our own decision-model service. Its two models run on our own GPUs in the
United States, and both are System One models: the state, the questions and the answers
above are their format too. Each reads English. What is their own:

| | Decider 2B | Kev 4B |
|---|---|---|
| Catalog id | `routerplus/decider-2b` | `routerplus/kev-4b` |
| Name at RouterPlus | `decider-2b` | `kev-4b` |
| Input, per million tokens | $0 | $0 |
| Context | 25,600 tokens; a longer input is cut (see below) | 8,192 tokens; a longer input is refused |

What holds for both:

| | Decider 2B and Kev 4B |
|---|---|
| Route | RouterPlus's endpoint only (`POST /v1/decisions` there, provider id `routerplus`). RouterPlus's short names in the table above are not ids here. One deployment serves both models, so one pool, one health circuit and one pause after a 429 cover both. |
| Price | Free: input and output are $0, and `usage.cost` is `0` on every call. RouterPlus is our own service, so the price is ours to set, not a provider's price passed through. See [Billing](#billing). |
| Questions | The gateway checks the System One rules above. RouterPlus takes every form that Jev takes, on both models: `instructions` as a string or as a JSON object or array, `score` levels as strings or as JSON objects (the answer's `legend` gives the levels as you sent them), and a `choice` option's description as a string, JSON or `null`. In probes on 2026-10-02, each model also answered a `choice` with one option, a `choice` of 255 options, a `score` of 2 and of 10 levels, a yes/no with described `true` and `false`, a JSON object or array as the state, and 200 questions in one request. A question that RouterPlus calls malformed is its 400, relayed with RouterPlus's reason for each question id (see [Errors](#errors)); it costs nothing. |
| Long input | Decider 2B reads up to 25,600 tokens and cuts a longer input: it keeps the question and its options first, then the start of the state. It drops the rest and does not return an error. Kev 4B never cuts: above 8,192 tokens it refuses the request with a 400 (see [Errors](#errors)). The gateway bills RouterPlus's `usage.input_tokens`, and RouterPlus counts there only the tokens the model can read: at most the model's context. In a test on 2026-10-02, a Decider 2B request of 30,053 tokens billed 25,600. RouterPlus can refuse a long input in place of cutting it (its `truncate` field), but the gateway drops that field (see [Parameters that do not apply](#parameters-that-do-not-apply)), so through the gateway a long input on Decider 2B is always cut. Put what matters at the start of the state, or use Kev 4B for an input up to 8,192 tokens and Decider 2B up to 25,600. |
| Response cache | RouterPlus keeps each answer, with its usage, for 600 s (10 minutes). The cache key is the RouterPlus API key, the model and the exact request body, byte for byte. The gateway sends every request with our one RouterPlus key, so an identical request from any of our buyers within 600 s can get the cached answer: the same `answers`, and the same input count in `usage`. Your response keeps its own `id`. RouterPlus marks it with `usage.cached: true` and the header `X-Cache: HIT` in its own response. A cached answer costs $0: the listing prices cached input at $0, and the gateway bills a cached answer's input tokens at that price. The gateway's response marks it with `usage.cached_input_tokens`, equal to `usage.input_tokens`. RouterPlus skips its cache for a request with `"cache": false` or a `Cache-Control: no-cache` header. The gateway forwards neither, so you cannot skip the cache through the gateway. |
| Request id | RouterPlus sends `X-Request-Id` (`req_…`) on every answer. The gateway keeps it with the request's record as the provider's id. Your response's `id` and its `x-request-id` header are the gateway's own. |
| Cold start | After a quiet period, the first request starts the model on a GPU. This takes 15 to 20 s. The gateway waits up to 60 s for a decision, so a cold start is a slow 200, not an error. |
| Our limit | The deployment's pool takes 11,400 requests a minute, 15,000,000 tokens a minute and 59 requests in flight at once. Each organization may use 75 % of each: 8,550 requests a minute, 11,250,000 tokens a minute, 44 at once. The 59 in flight is the burst limit: a pool has no limit per second, and 59 requests in flight at RouterPlus's typical 0.31 s a call carry about 190 a second, just above the pool's 11,400 a minute; with our test environment's 10 a second, that stays within RouterPlus's own limit (next row). Your organization's own limits bind first: by default 300 requests a minute, 1,000,000 tokens a minute and 8 in flight (see [Rate limits & spend caps](/docs/limits)). A burst above any of these is a 429 `rate_limit` from the gateway, with `retry-after` and the `x-tm-limit-*` headers, before anything reaches RouterPlus; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown` for the pause after a relayed RouterPlus 429, next row). |
| RouterPlus's limit | About 200 requests a second for the whole endpoint, shared with our test environment. A cached answer comes back sooner than a computed one, so a burst of repeated requests can still reach it. Above it RouterPlus answers 429. The gateway relays it as 429 `rate_limit` with `x-tm-upstream-status: 429`, and pauses the deployment's pool for that `retry-after`, at least 1 s and at most 60 s. A relayed 429 never counts against the deployment's health circuit. |
| Busy or failed | A 5xx from RouterPlus is a retryable `upstream_error` that counts against the deployment's health circuit: two in a row open it for 30 s for both models, and in that window every request but one probe gets 503 `gateway_error`, "all deployments cooling down". Retry after a few seconds, with backoff. A RouterPlus 401, 402 or 403 refuses our key, never yours: you get 503 `model_unavailable`, as from every System One provider. See [Errors](#errors). |
| Retention | RouterPlus writes no request logs. Its answer cache keeps each answer and its usage for 600 s, under a key made from our RouterPlus key and the request. The cache is all it keeps, so the deployment declares `prompt_logging: "none"`. Nothing is used for training. See [Data policy](/docs/data-policy). |

#### Perplexity Decider v1 27B

Perplexity's Decider v1 27B is a System One model: the state, the questions and the answers
above are its format too. The gateway calls Perplexity's own API. One thing sets it apart
from the rest of the family, and it changes the cost: it reads and bills the state once per
question. What is its own:

| | Perplexity Decider v1 27B |
|---|---|
| Catalog id | `perplexity/pplx-decider-v1-27b`. Perplexity's own name `pplx-decider-v1-27b` is not an id here. |
| Route | Perplexity's API only (`POST https://api.perplexity.ai/v1/decisions`, provider id `perplexity-decisions`), as `pplx-decider-v1-27b`. The provider id is not `perplexity`, which is an OpenRouter host tag. Perplexity answers in Jev's shape, with no envelope. The gateway keeps Perplexity's `x-request-id` with the request's record as the provider's id. |
| Price | $0.04 per million input tokens. Output is free. This is Perplexity's own price, with no markup. See [Billing](#billing). |
| The state, once per question | Perplexity runs one prompt for each question, and each prompt holds the whole state. It bills the state in each one. In a test on 2026-10-02, a state of 1,843 tokens billed 1,843 tokens with one question and 9,215 with five. So on this model one call with ten questions saves round trips, not tokens: it costs about what ten calls cost. The gateway holds the state once per question too (see Our limit below). The question ids are not billed. `output_tokens` is one per question, at $0. A `choice` with one option runs no prompt and costs nothing: a request of only such questions comes back with 0 tokens and costs $0. |
| Questions | 1 to 128 questions per call. Above 128 is Perplexity's 400, "Each request needs between 1 and 128 questions", relayed as an `upstream_error` with Perplexity's words; it costs nothing. A `choice` takes 1 to 255 options, as on Jev. A `score` takes 2 to 10 levels. A question id is any key the gateway takes. The gateway does not check the 128 (the Playground's Decision mode does, before a run). |
| Structured content | `instructions`, a `score` level and a `choice` option's description may be a JSON object or array, as on Jev, and an option's description may be `null`. A probe with a JSON object as `instructions`, JSON option descriptions and JSON score levels answered 200 (2026-10-02). |
| Context | 262,144 tokens for each question's prompt: the state and that one question. A request with several questions can bill more than that in all: a state of 52,348 tokens with six questions billed 314,088 tokens (2026-10-02). Above the limit, Perplexity refuses the request with a 400, "Input length (262144) exceeds or equals model's maximum context length (262144)". The gateway returns it as 400 `context_overflow`, with `metadata.retryable: false`; it costs nothing. Perplexity does not cut a long state. |
| Image parts | Decision models take text and JSON. An object with `"type": "image_url"`, anywhere in the state or in a question, is a 400 from the gateway before any money is reserved, for example `state[1]: an object with "type": "image_url" is an image part, and model "perplexity/pplx-decider-v1-27b" takes text and JSON only`. Other JSON, a field named `type` with another value included, goes as you sent it. |
| Numbers | Full precision, as Nimble's. On a `choice`, `confidence` is Jev's: `(p_max − 1/n) / (1 − 1/n)`, so a threshold that you tuned on Jev's choice confidence transfers. On a `score` it does not. Perplexity's score confidence is `1 − Σ pᵢ·|i − m| / ((1/L)·Σⱼ |j − m|)`, where `m` is the most likely level and `L` the number of levels; Jev measures the second sum about the central level, not about `m`. The two agree only when the most likely level is the central one. On one live answer (three levels at 0.756, 0.193 and 0.051), Perplexity's `confidence` was 0.7051, where Jev's formula gives 0.5577. So tune a score threshold on Perplexity Decider, or use the probabilities. `confidence` is not the chance that the answer is right. |
| Busy | Perplexity answers 504 when the model does not answer in about a minute; the body is an HTML page. The gateway returns a retryable 504 `upstream_error` that counts against the deployment's health circuit, like every 5xx: two in a row open it for 30 s, and in that window every request but one probe gets 503 `gateway_error`, "all deployments cooling down". The gateway's own wait for a decision (60 s) also ends in a 504. In Perplexity's tests a few hundred input tokens answered in under 2 s, and a prompt near the context in 23 s. A Perplexity 401, 402 or 403 refuses our key, never yours: you get 503 `model_unavailable`, as from every System One provider. See [Errors](#errors). |
| Our limit | The deployment's pool takes 540 requests a minute, at most 8 requests at a time and 15,000,000 tokens a minute. Each organization may use 75 % of each: 405 requests a minute, 6 at a time, 11,250,000 tokens a minute. Your organization's own limits bind first: by default 300 requests a minute, 1,000,000 tokens a minute and 8 in flight (see [Rate limits & spend caps](/docs/limits)). Tokens are held at the gateway's estimate while a request runs, with the state counted once per question, and corrected to Perplexity's count when it settles. So a long state with many questions can be larger than a token limit on its own: such a request is a 429 `rate_limit` with `x-tm-limit-kind: tpm` before anything reaches Perplexity, and it cannot pass as it is. Ask fewer questions per call, or send a shorter state. A burst above any limit is a 429 `rate_limit` from the gateway, with `retry-after` and the `x-tm-limit-*` headers; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown` for the pause after a relayed Perplexity 429, next row). |
| Perplexity's limit | Perplexity allows our organization 10 requests a second, over a rolling second, shared with our test environment. Our pools keep to that as a minute budget (600 a minute in all), but they do not count seconds: a burst inside one second can pass 10, and then Perplexity's 429 limits it. It also limits large bursts of tokens, and publishes no number for that. Above a limit Perplexity answers 429, "Request rate limit exceeded, please try again later.", with `Retry-After: 1`. The gateway relays it as 429 `rate_limit` with `x-tm-upstream-status: 429`, and pauses the deployment's pool for that `retry-after`, at least 1 s and at most 60 s. A relayed 429 never counts against the deployment's health circuit. |
| Retention | Perplexity says that it does not retain query data sent through its API and does not train on it: its API has zero day retention of prompt data by default. Its compute runs on AWS in North America. The deployment declares `prompt_logging: "none"`. See [Data policy](/docs/data-policy). |

#### Sage 1.2

Levanto's Sage is a System One model: the state, the questions and the answers above are its
format too. The gateway calls Levanto's System One API. Like Perplexity Decider, it reads and
bills the state once per question. What is its own:

| | Sage 1.2 |
|---|---|
| Catalog id | `levanto/sage-1.2`. Levanto's own names (`sage-latest`, `levanto-sage-v1.3`) are not ids here. |
| Route | Levanto's API only (`POST https://sage.levanto.ai/v1/systemone`, provider id `levanto`), as `sage-latest`. Levanto answers in Jev's shape. |
| Price | $0.05 per million input tokens. Output is free. This is Levanto's own price, with no markup. See [Billing](#billing). |
| The state, once per question | Levanto reads the state once for each question and bills it each time. So on this model one call with ten questions saves round trips, not tokens. The gateway holds the state once per question too. |
| Questions | A `choice` takes 2 to 120 options and a `score` 2 to 26 levels. Every question gets an answer, never `null`. Levanto refuses a request outside its limits; the gateway relays the refusal with Levanto's words, and it costs nothing. |
| Not offered | Levanto's own `/decide` format (`kind`, sort, tags), images, grounding and `reasoning`. A question with `kind` is a 400; `reasoning` and `latency_mode` are dropped and recorded. |
| Retention | The deployment declares `prompt_logging: "retained"`. See [Data policy](/docs/data-policy). |

### GLiNER schema

GLiNER-2.5-Decide is a small encoder model. It reads one text and fills in a schema: it
classifies the text against your labels, and it can extract entities, structured fields
and relations from it.

**State.** Text only: a non-empty string, or `{"kind": "text", "value": "..."}`. A list,
an image or any other state is a 400 that names the field. To judge a record, send it as
text, for example the JSON as a string.

**Schema.** GLiNER takes no `questions`. It takes `schema`, Fastino's own schema object.
The gateway checks it and forwards it as it is. `schema` has only these four keys; send at
least one, and each key you send must be non-empty (an empty list or object is a 400 that
names the key). A flat list in place of the object is a 400 (Fastino has deprecated it).

| Key | Form | Limits |
|---|---|---|
| `classifications` | A list of `{"task", "labels", "multi_label"?, "top_k"?, "cls_threshold"?}` | 1 to 50 tasks. `task` is 1 to 256 characters, unique in the schema, and not `__proto__`. `labels` is 1 to 100 unique non-empty strings. `multi_label` is a boolean. `top_k` is an integer, 1 or more. `cls_threshold` is a number from 0 to 1. Any other key is a 400 that names it. |
| `entities` | A list of entity names, or of `{"name", "description"?}` | 1 to 50 entities. Each name is a non-empty string, unique in the list. |
| `structures` | An object of structure name to a list of fields, each `"field::type::description"` | 1 to 50 structures. A name is 1 to 256 characters and not `__proto__`. 1 to 50 non-empty field strings per structure. |
| `relations` | A list of relation names, or an object of name to `{"description"?, "threshold"?}` | 1 to 50 unique non-empty names. `threshold` is a number from 0 to 1. A `head` or `tail` key is a 400. |

**Labels are plain strings.** A label cannot be an object with a description. To describe a
label, put the description in the label itself: `"credit: store credit"`. Descriptions
matter: on the same refund ticket, the labels `refund`, `credit`, `deny` chose `refund`,
and the same labels with descriptions chose `credit`. The answer carries the label text
exactly as you sent it.

**The task name is the question.** A classification has no `instructions` field. GLiNER
reads the task name, so a task name can be a question: a task named
`"Does the customer ask for money back?"` with the labels `["yes", "no"]` is a yes/no
question.

**Result names share one namespace.** GLiNER puts every result at the top level of its
answer: each task under its task name, each structure under its structure name, the
entities under `entities` and the relations under `relation_extraction`. Two results with
one name would overwrite each other at Fastino without an error. So the gateway refuses a
collision with a 400 that names both, for example a task named `entities` in a request that
also asks for entities.

```json
{
  "model": "fastino/gliner-2.5-decide",
  "state": "I was charged twice for my March invoice. Please refund the order from Acme Corp. Signed, Jane Doe.",
  "schema": {
    "classifications": [
      { "task": "action", "labels": ["refund", "credit", "deny"] },
      { "task": "issues", "labels": ["double_charge", "angry_customer", "fraud"], "multi_label": true }
    ],
    "entities": ["person", { "name": "organization", "description": "business or institution name" }]
  }
}
```

#### GLiNER options

Three optional top-level fields go to Fastino as they are. Leave them out to use Fastino's
defaults.

| Field | Values | What it does |
|---|---|---|
| `threshold` | A number from 0 to 1. Default 0.5. | The confidence a result needs to be returned. A task's own `cls_threshold` wins for that task. Lower values return more results; higher values return fewer, surer ones. |
| `include_confidence` | Boolean. Default `true`. | With `false`, a single-label answer is the bare label string, and a multi-label answer is a list of label strings. |
| `include_spans` | Boolean. Default `true`. | With `true`, each entity carries `start` and `end`: its character offsets in the text. |

A single-label task returns its top label only; `top_k` does not add more labels to the
answer. When no label reaches the threshold, the task's answer is `null`: set
`"cls_threshold": 0` on the task to always get the top label and its confidence. A task with `multi_label: true` returns every label at or above the threshold, so
`"cls_threshold": 0` returns every label with its confidence.

#### Store

Fastino keeps each inference unless the request says `store: false`. **The gateway always
sends `store: false`.** A `store` field that you send is dropped and recorded, never
forwarded. Our Fastino account also has Zero Data Retention on, so Fastino does not train on
your text. See [Data policy](/docs/data-policy).

#### Not offered yet

These parts of Fastino's API are not reachable through the gateway today:

- Fine-tuned GLiNER models (Fastino's training-job ids).
- Fastino's own batch and async routes.
- Several messages. The gateway sends Fastino one user message: the state. GLiNER does not
  read a system message.

### Parameters that do not apply

Decision models take no sampling or output controls. The gateway forwards no `temperature`,
`max_tokens`, `stream`, `seed`, `user` or `metadata` to them, and none offers prompt
caching, so `cache_control` has nothing to act on. RouterPlus's answer cache (see
[Decider 2B and Kev 4B](#decider-2b-and-kev-4b)) is not prompt caching: it
works by itself on a whole identical request, and `cache_control` does not reach it. A
top-level field that the model does not take is **dropped and recorded** (D8 §2), never
forwarded. RouterPlus's own `truncate` and `cache` fields are dropped in this way:

| Model | Top-level fields it takes | Examples of fields dropped |
|---|---|---|
| Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sage | `model`, `state`, `questions`, `provider` | `temperature`, `reasoning`, `latency_mode`, `images`, `request_id`, `truncate`, `cache` |
| GLiNER | `model`, `state`, `schema`, `provider`, `threshold`, `include_confidence`, `include_spans` | `store`, `temperature`, `messages` |

The dropped names are in the `x-tm-dropped-params` response header and on the attempt's
ledger row. To refuse instead, send `"provider": {"require_parameters": true}`: a would-be
drop is then a 400 before any money is reserved.

## Response

HTTP 200, `application/json`. Every decision model uses the same envelope: `id`, `object`,
`model`, `answers` and `usage`.

### System One answers

The shape is the same from both of Jev's providers, from Mercury Decide (with
`"model": "inception/mercury-decide"` and `"cost": 0`), from Nimble (with
`"model": "bespokelabs/nimble-v3"` and its numbers at full precision), from Clef and
Clef-flash (with `"model": "cloudflare/clef"` or `"model": "cloudflare/clef-flash"`, its
numbers to 4 decimals and `"output_tokens": 0`), from Decider 2B and Kev 4B (with their
catalog ids in `model`) and from Perplexity Decider (with
`"model": "perplexity/pplx-decider-v1-27b"`, its numbers at full precision and one output
token per question):

```json
{
  "id": "5d0c1c4e-3f7a-4c55-9d1b-2f0e7f6f2a10",
  "object": "decision",
  "model": "typesafe/jev-1.13",
  "answers": {
    "team": { "type": "choice", "choice": "billing",
              "probabilities": { "billing": 0.88, "technical": 0.12, "sales": 0 }, "confidence": 0.81 },
    "frustration": { "type": "score", "score": 1.05,
                     "legend": { "0": "Calm", "1": "Annoyed", "2": "Angry" },
                     "probabilities": { "0": 0, "1": 0.95, "2": 0.05 }, "confidence": 0.92 },
    "urgent": { "type": "noul", "noul": 0.95 }
  },
  "usage": { "input_tokens": 436, "output_tokens": 71, "cost": 0.000018 }
}
```

`answers` holds one typed answer per question, under the ids you sent, exactly as the
provider returned them.

### GLiNER answers

`answers` is GLiNER's own result, verbatim: the JSON that Fastino returns as a string in
`choices[0].message.content`, parsed into an object. The gateway adds nothing and removes
nothing. For the request in [GLiNER schema](#gliner-schema):

```json
{
  "id": "0b7e4c2a-6f1d-4a39-9c85-3e2d1f7a8b64",
  "object": "decision",
  "model": "fastino/gliner-2.5-decide",
  "answers": {
    "action": { "label": "refund", "confidence": 0.9327635765075684 },
    "issues": [ { "label": "double_charge", "confidence": 0.872347354888916 } ],
    "entities": {
      "person": [ { "text": "Jane Doe", "confidence": 0.99609375, "start": 90, "end": 98 } ],
      "organization": [ { "text": "Acme Corp", "confidence": 0.98828125, "start": 71, "end": 80 } ]
    }
  },
  "usage": { "input_tokens": 33, "output_tokens": 91, "cost": 0 }
}
```

| You asked | Key in `answers` | Value |
|---|---|---|
| A single-label task | The task name | `{"label", "confidence"}`: the top label. `null` when no label reaches the threshold (0.5 unless you set one; `"cls_threshold": 0` always returns the top label). With `include_confidence: false`, the bare label string. |
| A multi-label task | The task name | `[{"label", "confidence"}, ...]`: every label at or above the threshold, possibly none. With `include_confidence: false`, a list of label strings. |
| Entities | `entities` | An object of entity name to `[{"text", "confidence", "start", "end"}, ...]`. `start` and `end` are there when `include_spans` is on. |
| A structure | The structure name | A list of records, each field as `{"text", "confidence"}`, for example `{"order": [{"id": {"text": "4411", "confidence": 1.0}}]}`. |
| Relations | `relation_extraction` | `{"relations": [...]}`. |

- **A confidence is not a calibrated probability.** It is GLiNER's score for that label or
  span, from 0 to 1. Multi-label confidences are independent and do not sum to 1. Compare
  them within one model, not with System One's probabilities.
- GLiNER has no "not sure" answer. A low confidence is the sign that it is unsure.
- The gateway accepts a 200 only when the content parses to a JSON object with every task
  and structure name you asked, plus `entities` and `relation_extraction` when you asked for
  them. Anything else is an `upstream_error` and costs nothing.

### Usage and headers

| Field | Meaning |
|---|---|
| `id` | The gateway's request id, also in the `x-request-id` header. Look the request up with `GET /v1/generation?id=`. |
| `model` | The catalog id that was billed. The provider's own name for the model is in the `x-tm-upstream-model` header. |
| `usage.input_tokens` | Input tokens, billed at the model's input rate. For the System One models and GLiNER, the provider's count (Fastino's `prompt_tokens`); on Decider 2B, at most the model's context; on Perplexity Decider and Sage, the state once per question. |
| `usage.output_tokens` | Output tokens, billed at the model's output rate. Every decision model prices output at $0; Sage, Clef and Clef-flash report 0, and Perplexity Decider one per question. For GLiNER, Fastino's `completion_tokens`. |
| `usage.cached_input_tokens` | Present only on an answer from RouterPlus's cache (Decider 2B, Kev 4B). The part of `usage.input_tokens` that the cache answered: all of it. These tokens are billed at the cached-input price, $0. |
| `usage.cost` | USD, the full debit for the request, after any discount, including any attempt that failed over before this one. `0` on BYOK. Absent on static dev keys. |
| `usage.cost_before_discount` | Present only when a discount applied. What the request costs at the list price. |
| `usage.discount_percent` | Present only when a discount applied. The percent taken off. |

Response headers: `x-request-id`, `x-tm-provider` (`typesafe`, `openrouter-decisions`,
`openrouter-decisions-free`, `bespokelabs`, `workers-ai`, `routerplus`,
`perplexity-decisions`, `levanto` or `fastino`), `x-tm-served-by` (the provider OpenRouter names, when it names one: `Inception`
for Mercury Decide), `x-tm-upstream-model` (the provider's own name for the model, for
example `clef` or `clef-flash` from `workers-ai`, `decider-2b` from `routerplus`,
`pplx-decider-v1-27b` from `perplexity-decisions`),
`x-tm-attempts`, `x-tm-upstream-status`, `x-tm-discount-percent` when a discount applied
and, when something was dropped, `x-tm-dropped-params`.

## Billing

A decision is metered like any other request: key limits, the reservation, the ledger and
`usage.cost` all work as on the chat routes. Every decision model is priced per million
tokens. Jev, Nimble, Clef, Clef-flash, Perplexity Decider, Sage and GLiNER charge
for input only.
Mercury Decide is $0 both ways while OpenRouter serves only its free variant; a paid
variant, or Inception serving it directly, would be a new listed price, and this page would
say so. Decider 2B and Kev 4B run on RouterPlus, our own service, and they are free:
$0 in and out, so `usage.cost` is `0` on every call. Each charge rounds down to the micro-dollar. Perplexity Decider bills the state once
per question (see [Perplexity Decider v1 27B](#perplexity-decider-v1-27b)): its cost grows
with the number of questions as well as with the state.

| Model | Input, per million tokens | Output, per million tokens | What one call costs |
|---|---|---|---|
| Jev | $0.042 | $0 | A three-question call of about 400 input tokens costs about $0.000017. |
| Mercury Decide | $0 | $0 | Nothing. The call is still metered, limited and recorded in `GET /v1/generation`, and it still needs an organization with credit. |
| Nimble | $0.04 | $0 | A three-question call of about 340 input tokens costs $0.000013. Bespoke's own price, with no markup. |
| Clef | $0.24 | $0 | A call of 400 input tokens costs $0.000096. Cloudflare's own price, with no markup. |
| Clef-flash | $0.09 | $0 | A call of 400 input tokens costs $0.000036. Cloudflare's own price, with no markup. |
| Decider 2B | $0 | $0 | Nothing, as on Mercury Decide. |
| Kev 4B | $0 | $0 | Nothing, as on Mercury Decide. |
| Perplexity Decider | $0.04 | $0 | Five questions about a state of 1,843 tokens bill 9,215 input tokens: $0.000368. Perplexity's own price, with no markup. |
| Sage | $0.05 | $0 | Three questions about a short ticket read 204 input tokens: $0.00001. Levanto's own price, with no markup. |
| GLiNER | $0.03 | $0 | A call of 1,000 input tokens costs $0.00003. A call of 33 input tokens, as in the example in [GLiNER answers](#gliner-answers), rounds down to $0. |

If a provider answers 200 without a usage object, the gateway bills its own estimate, the
same amount it reserved, and marks the attempt `estimated`. For GLiNER it comes from the
size of the text and from the tasks, labels and extraction keys in the schema. For
Perplexity Decider and Sage it counts the state once per question. A 200 that
does not answer every question, or whose GLiNER content is
not the object described in [GLiNER answers](#gliner-answers), is an `upstream_error` and
costs nothing. After a provider answers 200, the gateway never sends the same request to a second
provider.

### Discount

A discount may apply to decision models: to every decision model, or to some of them, each
at its own percent. It applies from every key, including the playground's. The list price
does not change, and the discount can change or end: read it from each response, not from
this page. A discounted model's row in `/api/models.json` carries `discount_percent`, and its
page shows the percent beside the struck list price.

- `usage.cost` is what you pay, after the discount.
- `usage.cost_before_discount` is the list price of the request.
- `usage.discount_percent` and the `x-tm-discount-percent` header give the percent.
- When no discount applies, these three are absent.

A discounted charge is the list charge × (100 − percent) / 100, rounded down to the
micro-dollar. At 100 % it is 0, and the request is still metered, limited and recorded in
`GET /v1/generation`. The discount is for organizations with credit: a balance of $0 is
refused with `429 insufficient_quota` whatever the discount, so a new account must complete verified browser sign-in
and buy credits before its first decision.

On Mercury Decide, Decider 2B and Kev 4B the list price is already $0, so
the discount changes nothing: `usage.cost` and `usage.cost_before_discount` are both `0`.
When a discount covers one of them, the percent is still reported, as on every model it
covers, so your code can read one shape for all ten.

## Errors

Errors use the gateway's normal shape — see [Errors](/docs/errors).

| Status | `error.code` | When |
|---|---|---|
| 400 | `invalid_request` | The body failed a rule above: a question in another model's format, a field a question does not take. The message names the field, for example `questions.team.criteria: at most 255 options`. |
| 400 | `invalid_request` | The request does not fit the model's grammar: `schema` on Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider or Sage (the message says the model takes `questions`), or `questions` on GLiNER (the message says the model takes a GLiNER `schema`). |
| 400 | `invalid_request` | GLiNER only: a state that is not text, a flat-list schema, a schema key or classification key GLiNER does not take, a limit above, two tasks with one name, or two results with one name (the message names both). |
| 400 | `invalid_request` | The model is not a decision model. The message names the route that serves it. |
| 400 | `invalid_request` | Perplexity Decider: an object with `"type": "image_url"` anywhere in the state or in a question. Decision models take text and JSON. The message names where it is, for example `state[1]: an object with "type": "image_url" is an image part, and model "perplexity/pplx-decider-v1-27b" takes text and JSON only`. It is refused before any money is reserved. |
| 400 | `context_overflow` | Perplexity Decider: one question's prompt (the state and that question) is above 262,144 tokens. Perplexity's 400, relayed with its words, for example `upstream perplexity-decisions returned 400: Input length (262144) exceeds or equals model's maximum context length (262144)`. `metadata.retryable` is `false`: make the state shorter. It costs nothing. |
| 404 | `model_unavailable` | No decision model has this id. |
| 413 | `request_too_large` | The body is larger than 1 MB. |
| 413 | `context_overflow` | Clef and Clef-flash: Cloudflare estimated the request above the model's 65,536 tokens and refused it (its 413, code 5021). The message carries Cloudflare's words, for example `upstream workers-ai returned 413: The estimated number of input and maximum output tokens (130569) exceeded this model context window limit (65536).` `metadata.retryable` is `false`: make the state shorter. It costs nothing. |
| 503 | `model_unavailable` | Sage: Levanto refused our account (its 401, 402 or 403, for example when our month's usage is used up); the message says the model is out of capacity at the provider. GLiNER: Fastino refused our account for credit or billing (its 402 or 403). The message says the model is out of capacity at the provider. Nimble: Bespoke refused our account for credit (its 402); the message says the model is out of capacity at the provider, and `metadata.retryable` is `false`. Clef and Clef-flash: Cloudflare refused our account, for our token or plan (its 401 or 403), or because the account's free daily allocation is used up (its 429 with code 3036, until 00:00 UTC). The message says the model is out of capacity at the provider, and `metadata.retryable` is `false`. A 401, 402 or 403 from any System One provider (TypeSafe's or OpenRouter's, on Jev or Mercury Decide, and RouterPlus's, on Decider 2B and Kev 4B, and Perplexity's, on Perplexity Decider, too) reads the same: it is our account, never yours. On a connection with your own provider key, that provider's 401, 402 or 403 is your account, and it keeps its status. Other models are not affected. |
| 503 | `gateway_error` | "all deployments cooling down", with `retry-after: 5`: the deployment's health circuit is open. On Sage, GLiNER, Nimble, Clef, Clef-flash, the two RouterPlus models and Perplexity Decider this is also what most requests see while the provider refuses our account: each refusal counts against the provider's circuit, so after two of them only one request in each 30 s reaches the provider and gets the message above, and the rest get this one. Decider 2B and Kev 4B share one deployment and so one circuit: two RouterPlus 5xx in a row open it for both. A provider's 429 never counts, so Mercury Decide's daily cap on OpenRouter, Nimble's limit of 8 requests at once, Cloudflare's busy answer on Clef, RouterPlus's limit of about 200 a second and Perplexity's limit of 10 a second stay 429s (below), and Cloudflare's code 3036 stays the 503 above. |
| 422 and other 4xx | `upstream_error` | The provider refused the request as wrong. The message carries the provider's own reason, and `x-tm-upstream-status` its status. The gateway does not send a request the provider called wrong to another provider. A Fastino 404 (an unknown upstream model id: our misconfiguration, never yours) is a 502 `upstream_error`, with `x-tm-upstream-status: 404`. A Nimble request above Nimble's own limits (more than 64 questions, a `choice` with one option, or a question's prompt above 32,768 tokens with the state) is Bespoke's 422, relayed here with Bespoke's words; it costs nothing. A Clef or Clef-flash request outside Clef's own limits (more than 64 questions, a question id that does not match `^[A-Za-z0-9_.-]{1,100}$`, or a `choice` with one option) is Cloudflare's 422 (code 5012) or 400 (code 5006), relayed here with Cloudflare's words, for example `upstream workers-ai returned 422: Request body failed validation: questions: Dictionary should have at most 64 items after validation, not 65`; it costs nothing. A request that RouterPlus refuses (Decider 2B, Kev 4B) is RouterPlus's 400, relayed here with RouterPlus's reason and `x-tm-upstream-status: 400`; it costs nothing. RouterPlus refuses a malformed question (`invalid questions`, then the reason for each question id, for example `score needs criteria: a list of 2 to 10 levels`) and a Kev 4B request above 8,192 tokens (the question's id, then `ContextOverflow: branch too long`). A Perplexity Decider request with more than 128 questions is Perplexity's 400, relayed here with Perplexity's words (`Each request needs between 1 and 128 questions`); it costs nothing. |
| 429 | `rate_limit` | Mercury Decide: more than its pool allows (20 requests a minute in all, 15 per organization; 5 in flight at once, 3 per organization; 700,000 tokens a minute, 525,000 per organization), with `retry-after` — the gateway refuses before OpenRouter does. `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown`: the pool's pause after a relayed OpenRouter 429, 1 to 60 s) and `x-tm-limit-id` says whether it was the pool (`pool:…`) or your organization's share (`pool-share:…`). Or OpenRouter's daily cap on free requests is used up (1,000 a day across our account): OpenRouter's 429, relayed with its `retry-after` and `x-tm-upstream-status: 429`, until its daily reset; it never opens the deployment's circuit, and it pauses the pool for at most 60 s. |
| 429 | `rate_limit` | Nimble: more than its pool allows (7 in flight at once, 5 per organization; 600 requests a minute, 450 per organization; 15,000,000 tokens a minute, 11,250,000 per organization), with `retry-after`, before anything reaches Bespoke; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown`). Or Bespoke already runs 8 requests for our account: Bespoke's 429, relayed with its `Retry-After` and `x-tm-upstream-status: 429`; it never opens the deployment's circuit, and it pauses the pool for at most 60 s. |
| 429 | `rate_limit` | Clef and Clef-flash: more than their shared pool allows (200 requests a minute, 150 per organization; at most 8 requests at a time, 6 per organization; 2,000,000 tokens a minute, 1,500,000 per organization), with `retry-after`, before anything reaches Cloudflare; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown`). Or Workers AI is busy: Cloudflare's 429 with code 3040, relayed with Cloudflare's words (`upstream workers-ai returned 429: Capacity temporarily exceeded, please try again.`), its `Retry-After` when it sends one, and `x-tm-upstream-status: 429`. It never opens the deployment's circuit, and it pauses the pool for at most 60 s (1 s when Cloudflare sends no `Retry-After`). Any other Cloudflare 429 is relayed the same way, except code 3036 (the 503 above). |
| 429 | `rate_limit` | Decider 2B and Kev 4B: more than their shared pool allows (11,400 requests a minute, 8,550 per organization; 15,000,000 tokens a minute, 11,250,000 per organization; 59 in flight at once, 44 per organization), or more than your organization's own limits, which bind first (by default 300 requests a minute and 8 in flight), with `retry-after`, before anything reaches RouterPlus; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown`). Or RouterPlus is at its limit of about 200 requests a second for the whole endpoint: RouterPlus's 429, relayed with its `retry-after` and `x-tm-upstream-status: 429`; it never opens the deployment's circuit, and it pauses the pool for at most 60 s, for both models. |
| 429 | `rate_limit` | Perplexity Decider: more than its pool allows (540 requests a minute, 405 per organization; at most 8 requests at a time, 6 per organization; 15,000,000 tokens a minute, 11,250,000 per organization), or more than your organization's own limits, which bind first (by default 300 requests a minute, 1,000,000 tokens a minute and 8 in flight), with `retry-after`, before anything reaches Perplexity; `x-tm-limit-kind` names the limit (`rpm`, `tpm`, `concurrency`, or `cooldown`). The state is held once per question, so a long state with many questions can be larger than a token limit on its own (`x-tm-limit-kind: tpm`): ask fewer questions per call, or send a shorter state. Or Perplexity is at its limit of 10 requests a second for our organization: Perplexity's 429, relayed with its words (`Request rate limit exceeded, please try again later.`), its `Retry-After` and `x-tm-upstream-status: 429`; it never opens the deployment's circuit, and it pauses the pool for at most 60 s. |
| 504 | `upstream_error` | A System One provider answered 408, its own timeout. On Clef and Clef-flash the message is `upstream workers-ai returned 408: Request timeout`, with `x-tm-upstream-status: 408`. Retry after a few seconds. A 408 is not tried on a second provider, and it does not count against the deployment's circuit. The gateway's own wait for a decision (60 s) also ends in a 504 (next row). |
| 429, 5xx (and 529) | `rate_limit`, `upstream_error` | Every provider failed. A rate limit, an outage or an overload at one provider is tried on the next one first. Levanto answers 503 while Sage loads: retry after a few seconds. Fastino answers 425 or 503 while GLiNER starts (a cold start), and a cold start can also outlast the gateway's wait for a decision (60 s, a 504): both are a retryable `upstream_error`, so retry after a few seconds. A Fastino 429 is a `rate_limit`. Bespoke answers 503 with `Retry-After` while Nimble starts, 529 when it is busy and 502 when the model fails: each is a retryable `upstream_error` that counts against the deployment's circuit (two in a row open it for 30 s), so retry after a second or two. Cloudflare answers 500 when Clef fails ("Model execution failed"): a retryable `upstream_error` that counts against the circuit in the same way. A RouterPlus 5xx is a retryable `upstream_error` too, and it counts against the one circuit of Decider 2B and Kev 4B: two in a row open it for 30 s for both. A RouterPlus cold start (15 to 20 s) is not an error: it fits inside the gateway's 60 s wait and answers 200. Perplexity answers 504 when its model does not answer in about a minute (an HTML page, so the message has no words of Perplexity's): a retryable `upstream_error` that counts against the circuit in the same way. |
| 401, 402, 403 | `auth` | Every provider refused our own account with it. This is never your key's fault; it is tried on the next provider first. For GLiNER a 402 or 403, and for every System One model (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage) a 401, 402 or 403, is the 503 above. |

Jev reads at most 64 000 tokens per request through TypeSafe (32 000 for the state plus the
longest question), and 32 000 through OpenRouter. Mercury Decide reads at most 32 768 tokens
per request. Sage reads at most 32 768 tokens per request, the state and every question
together. Nimble reads at most 32,768 tokens for each question's prompt, the state included,
and never cuts a prompt. Clef and Clef-flash read at most 65,536 tokens per request, the
state and every question together, as Cloudflare estimates them (about 4 characters a
token); Cloudflare refuses a longer request with the 413 `context_overflow` above. Decider
2B's context is 25,600 tokens and Kev 4B's 8,192.
GLiNER-2.5-Decide's context is 8,192 tokens. A request above a provider's limit is refused
by that provider. Decider 2B cuts it instead, and bills only the tokens the model read; Kev
4B refuses it with a 400. See [Decider 2B and Kev 4B](#decider-2b-and-kev-4b).
Perplexity Decider reads at most 262,144 tokens for each question's prompt (the state and
that question), and never cuts one: Perplexity refuses a longer one with the 400
`context_overflow` above. See [Perplexity Decider v1 27B](#perplexity-decider-v1-27b).

## Examples

### Jev

```bash
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"typesafe/jev-1.13",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
```

Python — no SDK has a decisions method, so send plain HTTP:

```python
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "typesafe/jev-1.13",
        "state": "Help! My payouts have been failing for 3 days.",
        "questions": {"team": {"type": "choice", "instructions": "Which team?",
                               "criteria": {"billing": "Payments", "technical": "Bugs"}}},
    },
)
answer = r.json()["answers"]["team"]
print(answer["choice"], answer["confidence"])
```

Node:

```js
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "typesafe/jev-1.13",
    state: { ticket: "Help! My payouts have been failing for 3 days." },
    questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } },
  }),
});
const { answers, usage } = await r.json();
console.log(answers.urgent.noul, usage.cost);
```

### Mercury Decide

The same System One body as for Jev, under Mercury Decide's id. A 429 `rate_limit` here is
either our pool (wait `retry-after`) or OpenRouter's daily cap on free requests (wait for its
reset): switch on the class, as [Errors](/docs/errors) says, and read the message.

```bash
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"inception/mercury-decide",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
```

Python — plain HTTP; the answers read exactly as Jev's:

```python
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "inception/mercury-decide",
        "state": "Help! My payouts have been failing for 3 days.",
        "questions": {
            "team": {"type": "choice", "instructions": "Which team?",
                     "criteria": {"billing": "Payments", "technical": "Bugs"}},
            "angry": {"type": "score", "instructions": "How angry is the customer?",
                      "criteria": ["Calm", "Mild", "Moderate", "Upset", "Furious"]},
        },
    },
)
if r.status_code == 429:
    print("rate limited:", r.json()["error"]["message"], "retry after", r.headers.get("retry-after"))
else:
    answers = r.json()["answers"]
    print(answers["team"]["choice"], answers["angry"]["score"], r.json()["usage"]["cost"])  # cost is 0
```

Node:

```js
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "inception/mercury-decide",
    state: { ticket: "Help! My payouts have been failing for 3 days." },
    questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } },
  }),
});
const { answers, usage } = await r.json();
console.log(answers.urgent.noul, usage.cost); // 0.95, 0
```

### Nimble

The same System One body as for Jev, under Nimble's id. Keep to Nimble's limits: 1 to 64
questions, and 2 or more options in a choice.

```bash
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"bespokelabs/nimble-v3",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
```

Python — plain HTTP; the answers read exactly as Jev's, at full precision:

```python
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "bespokelabs/nimble-v3",
        "state": "Help! My payouts have been failing for 3 days.",
        "questions": {
            "team": {"type": "choice", "instructions": "Which team?",
                     "criteria": {"billing": "Payments", "technical": "Bugs"}},
            "urgent": {"type": "noul", "instructions": "Is this urgent?"},
        },
    },
)
answers = r.json()["answers"]
print(answers["team"]["choice"], answers["urgent"]["noul"])  # e.g. billing 0.997817283712868
```

Node:

```js
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "bespokelabs/nimble-v3",
    state: { ticket: "Help! My payouts have been failing for 3 days." },
    questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } },
  }),
});
const { answers, usage } = await r.json();
console.log(answers.urgent.noul, usage.input_tokens, usage.cost);
```

### Clef

The same System One body as for Jev, under Clef's or Clef-flash's id. Keep to Clef's limits:
1 to 64 questions, ids of 1 to 100 letters, digits, `_`, `.` or `-`, and 2 or more options
in a choice.

```bash
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"cloudflare/clef",
       "state":"Checkout has been failing for every customer for the last hour.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this support request urgent?"}}}'
```

Python — plain HTTP; the answers read as Jev's, to 4 decimals:

```python
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "cloudflare/clef",
        "state": "Checkout has been failing for every customer for the last hour.",
        "questions": {
            "urgent": {"type": "noul", "instructions": "Is this support request urgent?"},
            "team": {"type": "choice", "instructions": "Which team should handle this request?",
                     "criteria": {"billing": "Payments and invoices", "technical": "Bugs and outages",
                                  "sales": "Plans and upgrades"}},
            "severity": {"type": "score", "instructions": "How severe is the customer impact?",
                         "criteria": ["No impact", "Minor", "Major", "Critical"]},
        },
    },
)
answers = r.json()["answers"]
team = answers["team"]
print(team["choice"], team["probabilities"][team["choice"]], team["confidence"])  # e.g. technical 0.8274 0.5654
```

Node — Clef-flash, the smaller model:

```js
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "cloudflare/clef-flash",
    state: { ticket: "Checkout has been failing for every customer for the last hour." },
    questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } },
  }),
});
const { answers, usage } = await r.json();
console.log(answers.urgent.noul, usage.input_tokens, usage.output_tokens); // output_tokens is always 0
```

### RouterPlus models

The same System One body as for Jev, under Decider 2B's id. For Kev 4B, change only
`model`. A first request after a quiet period can take 15 to 20 s while the model starts,
so give your client a timeout above 60 s, the gateway's own wait.

```bash
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"routerplus/decider-2b",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
```

Python — plain HTTP, with a timeout that covers a cold start:

```python
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "routerplus/decider-2b",
        "state": "Help! My payouts have been failing for 3 days.",
        "questions": {
            "team": {"type": "choice", "instructions": "Which team?",
                     "criteria": {"billing": "Payments", "technical": "Bugs"}},
            "urgent": {"type": "noul", "instructions": "Is this urgent?"},
        },
    },
    timeout=70,
)
answers = r.json()["answers"]
print(answers["team"]["choice"], answers["urgent"]["noul"], r.headers["x-tm-upstream-model"])  # e.g. billing 0.97 decider-2b
```

Node:

```js
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "routerplus/decider-2b",
    state: { ticket: "Help! My payouts have been failing for 3 days." },
    questions: { urgent: { type: "noul", instructions: "Is `ticket` urgent?" } },
  }),
  signal: AbortSignal.timeout(70_000),
});
const { answers, usage } = await r.json();
console.log(answers.urgent.noul, usage.input_tokens, usage.cost);
```

### Perplexity

The same System One body as for Jev, under Perplexity Decider's id. Keep to its limits: 1 to
128 questions, and text or JSON only. Each question is billed with the whole state, so ask
only the questions you need.

```bash
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"perplexity/pplx-decider-v1-27b",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
```

Python — plain HTTP; the answers read as Jev's, at full precision:

```python
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "perplexity/pplx-decider-v1-27b",
        "state": "Help! My payouts have been failing for 3 days.",
        "questions": {"urgent": {"type": "noul", "instructions": "Is this urgent?"}},
    },
)
body = r.json()
print(body["answers"]["urgent"]["noul"], body["usage"]["input_tokens"])  # e.g. 0.9297849172276498 95
```

Node:

```js
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "perplexity/pplx-decider-v1-27b",
    state: { ticket: "Help! My payouts have been failing for 3 days." },
    questions: {
      team: { type: "choice", instructions: "Which team should handle `ticket`?", criteria: { billing: "Payments", technical: "Bugs" } },
      urgent: { type: "noul", instructions: "Is `ticket` urgent?" },
    },
  }),
});
const { answers, usage } = await r.json();
// Two questions: Perplexity bills the state twice.
console.log(answers.team.choice, answers.urgent.noul, usage.input_tokens, usage.cost);
```

### Sage

The same System One body as for Jev, under Sage's id. Each question is billed with the whole
state, so ask only the questions you need.

```bash
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"levanto/sage-1.2",
       "state":"Help! My payouts have been failing for 3 days.",
       "questions":{"urgent":{"type":"noul","instructions":"Is this urgent?"}}}'
```

### GLiNER

```bash
curl https://api.routerplus.com/v1/decisions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"fastino/gliner-2.5-decide",
       "state":"Help! My payouts have been failing for 3 days.",
       "schema":{"classifications":[{"task":"Is this urgent?","labels":["yes","no"]}]}}'
```

Python — plain HTTP; each answer sits under its task name:

```python
import os, httpx

r = httpx.post(
    "https://api.routerplus.com/v1/decisions",
    headers={"Authorization": f"Bearer {os.environ['TM_API_KEY']}"},
    json={
        "model": "fastino/gliner-2.5-decide",
        "state": "Help! My payouts have been failing for 3 days.",
        "schema": {"classifications": [{"task": "team",
                                        "labels": ["billing: payments", "technical: bugs"]}]},
    },
)
answer = r.json()["answers"]["team"]
print(answer["label"], answer["confidence"])  # a confidence, not a calibrated probability
```

Node:

```js
const r = await fetch("https://api.routerplus.com/v1/decisions", {
  method: "POST",
  headers: { authorization: `Bearer ${process.env.TM_API_KEY}`, "content-type": "application/json" },
  body: JSON.stringify({
    model: "fastino/gliner-2.5-decide",
    state: "Help! My payouts have been failing for 3 days.",
    schema: {
      classifications: [
        { task: "topics", labels: ["billing", "outage", "fraud"], multi_label: true, cls_threshold: 0 },
      ],
    },
  }),
});
const { answers, usage } = await r.json();
console.log(answers.topics, usage.input_tokens, usage.output_tokens, usage.cost);
```

## Model discovery

```bash
curl -s "https://api.routerplus.com/v1/models?output_modalities=decisions" \
  -H "Authorization: Bearer $TM_API_KEY"
```

Decision models carry `architecture.output_modalities: ["decisions"]` in `GET /v1/models`,
and `output_modalities: ["decisions"]` in the public feed `https://app.routerplus.com/api/models.json`.
The Anthropic shape of `GET /v1/models` never lists them.
