# Models & catalog

The catalog is one flat namespace of model ids. Every **chat** model in it is
callable through both wire surfaces — OpenAI-compatible `/v1/chat/completions`
and Anthropic-compatible `/v1/messages` — with one key; the gateway translates.
Image models answer on `POST /v1/images/generations` only, video models on
`POST /v1/videos` only, and decision models on `POST /v1/decisions` only.
Pricing is pass-through: the per-token rates in the catalog are what you pay,
and every billed response carries `usage.cost` so you can recompute the bill
from the wire.

There are four ways to read the catalog:

| Surface | URL | Auth | For |
|---|---|---|---|
| Model directory | `https://app.routerplus.com/` | none | Browsing, filters, 24-hour request counts |
| Per-model pages | `https://app.routerplus.com/models/<id>` | none | Providers with prices, latency, uptime and data policy; copy-paste snippets |
| Machine-readable JSON | `https://app.routerplus.com/api/models.json` | none | Scripts, dashboards, price checks |
| SDK-shaped list | `https://api.routerplus.com/v1/models` | API key | `client.models.list()` — see [GET /v1/models](/docs/api-models) |

## Model ids

Ids are exact strings. There is no fuzzy matching and no silent substitution: a
request for an id that is not in the catalog is a 404 that echoes the id you
sent (see [GET /v1/models](/docs/api-models)). The catalog as of this writing —
the live list is always [`/api/models.json`](https://app.routerplus.com/api/models.json).
Models served by two providers list both, first-party first — see
[One model, several providers](#one-model-several-providers):

| Model id | Providers | Context | Input /1M | Output /1M |
|---|---|---|---|---|
| `claude-sonnet-5-5` | anthropic, openrouter | 1M | $2 | $10 |
| `claude-opus-5-5` | anthropic, openrouter | 1M | $4 | $20 |
| `claude-fable-5-1` | anthropic, openrouter | 1M | $10 | $50 |
| `gpt-6.1-sol` | openrouter | 1.05M | $2 | $10 |
| `gpt-6-sol` | openrouter | 1.05M | $2 | $10 |
| `gpt-6-luna` | openrouter | 1.05M | $0.1 | $0.5 |
| `gpt-6-astra` | openrouter | 1.05M | $10 | $50 |
| `claude-sonnet-5` | anthropic, openrouter | 1M | $2 | $10 |
| `claude-opus-5` | anthropic, openrouter | 1M | $5 | $25 |
| `claude-fable-5` | anthropic, openrouter | 1M | $10 | $50 |
| `claude-sonnet-4-5` | anthropic, openrouter | 200K | $3 | $15 |
| `claude-haiku-4-5` | anthropic, openrouter | 200K | $1 | $5 |
| `gpt-4o-mini` | openai, openrouter | 128K | $0.15 | $0.6 |
| `gpt-4.1-mini` | openai, openrouter | 1.05M | $0.4 | $1.6 |
| `gpt-4o` | openai, openrouter | 128K | $2.5 | $10 |
| `tencent/hy4-preview` | openrouter | 1.05M | $0.834 | $2.501 |
| `z-ai/glm-5.3-flash` | openrouter | 1.31M | $0.075 | $0.25 |
| `deepseek/deepseek-v4.1-flash` | openrouter | 1.05M | $0.3 | $1.2 |
| `deepseek/deepseek-v4-flash-0731` | openrouter | 1.31M | $0.04998 | $0.09996 |
| `deepseek/deepseek-v4-flash` | openrouter | 1.05M | $0.08078 | $0.16156 |
| `tencent/hy3` | openrouter | 262K | $0.132 | $0.528 |
| `z-ai/glm-5.3` | openrouter | 1.31M | $1.4 | $4.4 |
| `xiaomi/mimo-v2.5` | openrouter | 1.05M | $0.14 | $0.28 |
| `z-ai/glm-5.2` | openrouter | 1.05M | $0.966 | $3.036 |
| `google/gemini-3.7-flash` | openrouter | 1.05M | $0.75 | $3.75 |
| `moonshotai/kimi-k3` | openrouter | 1.05M | $3 | $15 |
| `minimax/minimax-m3` | openrouter | 1.05M | $0.3 | $1.2 |
| `deepseek/deepseek-v4-pro` | openrouter | 1.05M | $0.687648 | $1.375296 |
| `upstage/solar-pro4` | openrouter | 524K | $0.03 | $0.12 |
| `deepseek/deepseek-v4-pro-0813` | openrouter | 1.05M | $1.12068 | $3.36204 |
| `google/gemini-3.8-flash` | openrouter | 1.05M | $0.75 | $3.75 |

A model or mix you deploy from [Optimize](/docs/model-search) gets an id of its
own, `tm/<name>-v<n>`. It is callable on `POST /v1/chat/completions` with a key of
the organization that owns it, and it is not listed in the catalog.

### Image models

Image models bill in tokens like everything else, but they answer on
[POST /v1/images/generations](/docs/api-images) and nowhere else. The per-image
ceiling is the most output tokens one image can bill; the reservation holds `n`
times it.

| Model id | Sizes | Qualities | Per-image ceiling | Text in /1M | Image out /1M |
|---|---|---|---|---|---|
| `gpt-image-1` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high` | 6,240 | $5 | $40 |
| `gpt-image-1-mini` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high` | 6,500 | $2 | $8 |
| `gpt-image-1.5` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high` | 6,500 | $5 | $32 |
| `gpt-image-2` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high` | 8,232 | $5 | $30 |
| `gpt-image-2.5-flare` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high · xhigh · max` | 8,232 | $5 | $30 |
| `gpt-image-2.5-sunburst` | `1024x1024 · 1536x1024 · 1024x1536` | `low · medium · high · xhigh · max` | 8,232 | $5 | $30 |

`size` and `quality` also accept `auto`, which lets the provider choose. Every
model's full descriptor map is in `/api/models.json` under
`supported_parameters`.

The Nano Banana, Seedream and Grok Imagine models are served through OpenRouter and billed
at the cost it reports ([how](/docs/pricing#models-billed-at-the-providers-reported-cost)).
They take a shape as `aspect_ratio` (`16:9`, `9:16`, …) instead of a `size`, most take a
`resolution` (`1K`, `2K`, `4K`), and Seedream takes a `seed`.

### Video models

Video models answer on [POST /v1/videos](/docs/api-videos) and nowhere else. A video is a
job: it is created at once and made in the next seconds to minutes. All three are served
through OpenRouter today and billed at the cost it reports.

| Model id | Made by | Length | Resolutions | Listed price | Per second ≤ |
|---|---|---|---|---|---|
| `seedance-2.5` | ByteDance | 4–30 s | 480p · 720p | $10.70 / M video tokens | $0.30 |
| `hailuo-3-max` | MiniMax | 5–15 s | 480p · 768p | $0.05–$0.08 per second | $0.08 |
| `wan-3.0` | Alibaba | 5–30 s | 480p · 720p · 1080p | $0.05–$0.20 per second | $0.20 |

### Decision models

Decision models answer on [POST /v1/decisions](/docs/api-decisions) and nowhere else. They
write no text: they return one typed answer per question about a state. Each model takes its
own question format.

| Model id | Made by | Providers | Question kinds | Input /1M |
|---|---|---|---|---|
| `typesafe/jev-1.13` | TypeSafe | `typesafe`, then `openrouter-decisions` | choice, score, yes/no (System One) | $0.042 |
| `inception/mercury-decide` | Inception | `openrouter-decisions-free` | choice, score, yes/no (System One, the same as Jev) | $0 |
| `bespokelabs/nimble-v3` | Bespoke Labs | `bespokelabs` | choice, score, yes/no (System One, the same as Jev) | $0.04 |
| `cloudflare/clef` | Cloudflare | `workers-ai` | choice, score, yes/no (System One, the same as Jev) | $0.24 |
| `cloudflare/clef-flash` | Cloudflare | `workers-ai` | choice, score, yes/no (System One, the same as Jev) | $0.09 |
| `routerplus/decider-2b` | RouterPlus | `routerplus` | choice, score, yes/no (System One, the same as Jev) | $0 |
| `routerplus/kev-4b` | RouterPlus | `routerplus` | choice, score, yes/no (System One, the same as Jev) | $0 |
| `perplexity/pplx-decider-v1-27b` | Perplexity | `perplexity-decisions` | choice, score, yes/no (System One, the same as Jev) | $0.04 |
| `levanto/sage-1.2` | Levanto | `levanto` | choice, score, yes/no (System One, the same as Jev) | $0.05 |
| `fastino/gliner-2.5-decide` | Fastino | `fastino` | classifications (single- and multi-label), entities, structures, relations | $0.03 |

Output is free on every decision model. Sage, like Perplexity Decider, bills the state once
per question. Mercury Decide is $0 in and out while
OpenRouter serves only its free variant; its pool takes 20 requests a minute (15 per
organization) and 5 in flight at once (3 per organization), and OpenRouter caps free
requests per day. Bespoke Nimble v3 is $0.04 per million input tokens, Bespoke's own price
with no markup. It takes 1 to 64 questions per call, and its pool takes 7 requests at once
(5 per organization): Bespoke allows our account 8, and one is kept for our test
environment. Clef and Clef-flash are Cloudflare's own decision models, served by Cloudflare
Workers AI at Cloudflare's own price with no markup: $0.24 and $0.09 per million input
tokens. They take 1 to 64 questions per call, each question id 1 to 100 letters, digits,
`_`, `.` or `-`, and a context of 65,536 tokens. They share one pool: 200 requests a minute
(150 per organization) and at most 8 at a time (6 per organization). Decider 2B and Kev 4B
are RouterPlus's own models, on our own GPUs in the United States, and free: $0 in and out.
They share one pool. Their contexts are 25,600 and 8,192 tokens.
RouterPlus answers an identical request from its cache for 600 s. Perplexity Decider v1 27B is Perplexity's own decision model, served by Perplexity's API
at Perplexity's own price with no markup: $0.04 per million input tokens. Perplexity bills
the state once per question, so a call costs about one state for each question it asks. It
takes 1 to 128 questions per call, and 262,144 tokens for each question's prompt. Its pool
takes 540 requests a minute (405 per organization) and 8 at once (6 per organization).
A discount may apply to decision models; see [Pricing](/docs/pricing).

Most models also carry cache pricing (`cached_prompt`, and `cache_write` where
the provider bills it) — the full per-SKU table is on each model's page and in
`/api/models.json`. A model with no separate cache or reasoning price bills
those tokens **at the base rate** — a missing SKU never means free.

### One model, several providers

A model id can be served by more than one provider — a first-party lab and an
aggregator, say. The catalog shows one card; the model page lists every
provider with its own uptime and TTFB. Routing tries providers in priority
order and fails over **before the first byte** only — never mid-answer.
Prices are identical across providers of one model (enforced when the catalog
is built), so which provider served you never changes your bill. The
`x-tm-provider` header says who did; when that provider knows the model under
its own id (an aggregator's `anthropic/claude-sonnet-4.5` for our
`claude-sonnet-4-5`), `x-tm-upstream-model` carries the id that was sent and
the response's `model` field echoes the provider's id truthfully. When an
aggregator names the provider it used, `x-tm-served-by` carries that name (for
example `Amazon Bedrock`).

### Date-pinned ids

The catalog lists **one clean id per model** (`claude-haiku-4-5`), but vendors
also ship dated snapshots (`claude-haiku-4-5-20251001`, `gpt-4o-2024-08-06`)
and `-latest` spellings, and tools like Claude Code send them. Any such pin of a
catalog model works without being listed: the gateway resolves it to its family
**for routing and pricing only**, and forwards your original id to the provider
verbatim — a snapshot pin is never silently retargeted to a newer version. The
response's `model` field shows the exact snapshot that served you. If a vendor
retires a pinned snapshot, you get the vendor's own error, not a quiet upgrade.
Unknown ids still 404 naming themselves, dated or not. A key bound to your own
provider connection matches ids literally, so list the exact ids you send there.

### Id grammar

Ids match `[a-zA-Z0-9][a-zA-Z0-9._/-]{0,127}` — letters, digits, `.` `_` `/`
`-`, max 128 chars. `:` is deliberately excluded: the suffix namespace
(`:something`) is reserved for future marketplace variants, so a provider can
never squat a routing suffix.

## Providers are manifests

A provider is a JSON manifest in the repo — endpoint, wire dialect, settlement
terms, data policy, and a model list with exact prices:

```json
{
  "provider": { "id": "anthropic", "name": "Anthropic (direct API)", "prompt_logging": "retained" },
  "endpoint": { "base_url": "https://api.anthropic.com/v1", "dialect": "anthropic" },
  "models": [{
    "id": "claude-haiku-4-5",
    "context_length": 200000,
    "max_output_tokens": 64000,
    "pricing": [
      { "type": "prompt", "unit": "token", "cost_usd_per_million": "1" },
      { "type": "completion", "unit": "token", "cost_usd_per_million": "5" }
    ]
  }]
}
```

Manifests compile into the routing catalog through a gated pipeline —
validate, emit a staging catalog, run the live conformance suite, then promote.
Adding or repricing a model is a **data change with a gate**, never a gateway
code deploy. Three lifecycle rules do real work:

- **`is_ready: false`** stages a model: validated and testable, but never
  routed and never listed.
- **`deprecation_date`** delists automatically: past that date the model drops
  out of the compiled catalog and the public pages.
- **Shared ids must agree.** If two providers declare the same model id with
  different prices, or one as a chat model and the other as an image model, the
  catalog build fails. Pass-through pricing stays coherent per id.

> [!NOTE]
> Prices sync as effective-dated append-only rows. A price change never
> rewrites history — old requests stay priced as dispatched.

## Conformance before listing

No model is listed on a provider's say-so. The conformance suite runs live
requests against the provider's endpoint, and passing it is the gate between
staging and live. It is deliberately implemented independently of the gateway's
own wire parsers, so a shared bug can't hide real breakage.

| Check | Proves |
|---|---|
| C1 | A non-stream completion returns real usage tokens (usage is the billable record) |
| C2 | Stream frames are parseable SSE |
| C3 | Streams terminate properly (`[DONE]` / `message_stop`) |
| C4 | Usage tokens are reported in-stream |
| C5 | An unknown model yields a parseable 4xx — not a silent fallback |
| C6 | Responses carry actual assistant content, not just billable counters |
| C7 | The response echoes the requested model id — no silent substitution |
| I1 | An image generation at the top declared quality returns base64 data and usage |
| I2 | The declared per-image ceiling covers the billed output |
| I3 | An undeclared size is refused |
| I4 | The ceiling holds at every declared size (release-day probe, opt-in: it buys a real render per size) |
| V1 | A video job at the cheapest tier finishes, reports its cost and downloads as an MP4 |
| V2 | The declared per-second ceiling covers the reported cost |
| V3 | An undeclared shape is refused |
| V4 | The ceiling holds at the top resolution (opt-in: it buys one more job per model) |

Image listings run I1–I3 and C5; video listings run V1–V3 and C5. C2–C4, C6 and
C7 do not apply to them — the media routes do not stream, carry no assistant
text, and an images response has no `model` field to echo.

C6 and C7 exist because the failure modes that matter are billing for an empty
answer and quietly serving a cheaper model. C7 accepts exactly one alias form:
a dated snapshot of the *same* id (`gpt-4o` → `gpt-4o-2024-08-06`), or an alias
target the manifest declares up front in `resolves_to`. `gpt-4o` answered by
`gpt-4o-mini` fails.

## The directory

[`https://app.routerplus.com/`](https://app.routerplus.com/) lists every live model in three sections —
chat, image and video — because they are called on different routes. A row shows
the model's context length and max output (for a chat model), or its sizes,
qualities, lengths and resolutions (for an image or video model), plus its
24-hour request count. The list can be searched (`/?q=haiku` matches model,
lab and provider names), filtered by kind, context length, supported parameters,
zero data retention, input price, lab and provider, and sorted by use, name,
price or context. Prices, latency and uptime depend on the provider, so they
live on each model's page at `/models/<id>`: one row per provider that runs the
model, per-SKU pricing, the provider's prompt-retention policy (marketplace
logs are content-free either way), and ready-to-paste snippets.

### The uptime figure is allowed to say nothing

Uptime and latency come from the trailing 24 hours of real routed traffic.
When a provider we call directly has fewer than 100 counted attempts in that
window, the model page shows **`n<100`** instead of a percentage. For a
provider reached through an aggregator, the figure is the one the aggregator
publishes.

> [!NOTE]
> A model with 3 requests is not "100% up", and we won't render it that way.
> No number until the sample is real — that rule is load-bearing, not a
> placeholder.

The hourly bars on a model page are green at ≥ 99%, amber at ≥ 95%, and red
below that; an hour with no sample stays grey.

## Machine-readable: /api/models.json

Everything above, as JSON, no auth:

```bash
curl -s https://app.routerplus.com/api/models.json
```

```json
{
  "models": [
    {
      "id": "claude-haiku-4-5",
      "display_name": "Claude Haiku 4.5",
      "output_modalities": ["text"],
      "providers": [
        { "id": "anthropic", "name": "Anthropic (direct API)", "dialect": "anthropic", "prompt_logging": "retained" },
        { "id": "openrouter", "name": "OpenRouter", "dialect": "openai", "prompt_logging": "retained",
          "upstream_id": "anthropic/claude-haiku-4.5" }
      ],
      "served_by": [
        { "id": "anthropic", "name": "Anthropic", "zdr": false,
          "routes": [ { "via": "anthropic", "direct": true }, { "via": "openrouter", "direct": false } ] },
        { "id": "amazon", "name": "Amazon", "zdr": true, "routes": [ { "via": "openrouter", "direct": false } ] },
        { "id": "azure", "name": "Azure", "zdr": false, "routes": [ { "via": "openrouter", "direct": false } ] },
        { "id": "google-vertex", "name": "Google Vertex", "zdr": true, "routes": [ { "via": "openrouter", "direct": false } ] }
      ],
      "context_length": 200000,
      "max_output_tokens": 64000,
      "pricing_usd_per_million": {
        "prompt": "1", "cached_prompt": "0.10",
        "cache_write": "1.25", "completion": "5"
      },
      "lab": { "id": "anthropic", "name": "Anthropic" },
      "weights": "proprietary"
    }
  ],
  "gateway": "https://api.routerplus.com/v1",
  "docs": "https://app.routerplus.com/docs"
}
```

| Field | Meaning |
|---|---|
| `id` | The exact string to put in your request's `model` field |
| `output_modalities` | `["text"]`, `["image"]`, `["video"]` or `["decisions"]` — which route serves this model. Filter on it; never guess from the id |
| `providers` | Every deployment serving this id, primary first, each with `id`, `name`, `dialect`, `prompt_logging`, and `upstream_id` when the provider knows the model under its own name |
| `dialect` (per provider) | The provider's **native** wire format — informational; every chat model answers on both surfaces |
| `prompt_logging` (per provider) | Provider's prompt retention: `none` or `retained` |
| `served_by` | The providers that run the model, one entry per provider however it is reached. `routes[].via` is the `providers` entry the request goes to; `direct: false` means an aggregator picks this provider per request. `quantizations` appears when the aggregator publishes it (for example `["fp8"]`). `zdr` is `true` when the provider offers Zero Data Retention for this model |
| `context_length` / `max_output_tokens` | Token limits; `max_output_tokens` may be `null`, and `context_length` is `null` for image and video models |
| `supported_parameters` | Image and video models only: the descriptor map the route enforces (`size`, `quality`, `n`, `seconds`, and the rest) |
| `pricing_usd_per_million` | Decimal-string USD per 1M tokens, keyed by SKU type |
| `billing` | `"reported_cost"` on the models billed at the provider's reported cost; absent on token-billed models |
| `price_card` | On those models, the provider's own listed prices, for reading — see [Pricing & billing](/docs/pricing#models-billed-at-the-providers-reported-cost) |
| `lab` | Who made the model, as distinct from who serves it — `null` when unknown |
| `weights` | `"open"` when the lab publishes the model's weights, `"proprietary"` when it never has, `null` when we are not certain |

For an image model `max_output_tokens` is the **per-image ceiling**, not a cap
on a completion. For a video model it is the ceiling per second of video, in
cost units (one micro-dollar each).

Prices are decimal **strings** so nothing is lost to float rounding — parse
them with a decimal type if you're doing money math. For the authenticated,
SDK-shaped list (what `client.models.list()` calls), see
[GET /v1/models](/docs/api-models).
