Models & catalog
The catalog, model ids, uptime honesty, machine-readable feeds.
The catalog is one flat namespace of model ids. Every chat model in it is callable through both wire surfaces — OpenAI-compatible /v1/chat/completions and Anthropic-compatible /v1/messages — with one key; the gateway translates. Image models answer on POST /v1/images/generations only, video models on POST /v1/videos only, and decision models on POST /v1/decisions only. Pricing is pass-through: the per-token rates in the catalog are what you pay, and every billed response carries usage.cost so you can recompute the bill from the wire.
There are four ways to read the catalog:
| Surface | URL | Auth | For |
|---|---|---|---|
| Model directory | https://app.routerplus.com/ | none | Browsing, filters, 24-hour request counts |
| Per-model pages | https://app.routerplus.com/models/<id> | none | Providers with prices, latency, uptime and data policy; copy-paste snippets |
| Machine-readable JSON | https://app.routerplus.com/api/models.json | none | Scripts, dashboards, price checks |
| SDK-shaped list | https://api.routerplus.com/v1/models | API key | client.models.list() — see GET /v1/models |
Model ids
Ids are exact strings. There is no fuzzy matching and no silent substitution: a request for an id that is not in the catalog is a 404 that echoes the id you sent (see GET /v1/models). The catalog as of this writing — the live list is always /api/models.json. Models served by two providers list both, first-party first — see One model, several providers:
| Model id | Providers | Context | Input /1M | Output /1M |
|---|---|---|---|---|
claude-sonnet-5-5 | anthropic, openrouter | 1M | $2 | $10 |
claude-opus-5-5 | anthropic, openrouter | 1M | $4 | $20 |
claude-fable-5-1 | anthropic, openrouter | 1M | $10 | $50 |
gpt-6.1-sol | openrouter | 1.05M | $2 | $10 |
gpt-6-sol | openrouter | 1.05M | $2 | $10 |
gpt-6-luna | openrouter | 1.05M | $0.1 | $0.5 |
gpt-6-astra | openrouter | 1.05M | $10 | $50 |
claude-sonnet-5 | anthropic, openrouter | 1M | $2 | $10 |
claude-opus-5 | anthropic, openrouter | 1M | $5 | $25 |
claude-fable-5 | anthropic, openrouter | 1M | $10 | $50 |
claude-sonnet-4-5 | anthropic, openrouter | 200K | $3 | $15 |
claude-haiku-4-5 | anthropic, openrouter | 200K | $1 | $5 |
gpt-4o-mini | openai, openrouter | 128K | $0.15 | $0.6 |
gpt-4.1-mini | openai, openrouter | 1.05M | $0.4 | $1.6 |
gpt-4o | openai, openrouter | 128K | $2.5 | $10 |
tencent/hy4-preview | openrouter | 1.05M | $0.834 | $2.501 |
z-ai/glm-5.3-flash | openrouter | 1.31M | $0.075 | $0.25 |
deepseek/deepseek-v4.1-flash | openrouter | 1.05M | $0.3 | $1.2 |
deepseek/deepseek-v4-flash-0731 | openrouter | 1.31M | $0.04998 | $0.09996 |
deepseek/deepseek-v4-flash | openrouter | 1.05M | $0.08078 | $0.16156 |
tencent/hy3 | openrouter | 262K | $0.132 | $0.528 |
z-ai/glm-5.3 | openrouter | 1.31M | $1.4 | $4.4 |
xiaomi/mimo-v2.5 | openrouter | 1.05M | $0.14 | $0.28 |
z-ai/glm-5.2 | openrouter | 1.05M | $0.966 | $3.036 |
google/gemini-3.7-flash | openrouter | 1.05M | $0.75 | $3.75 |
moonshotai/kimi-k3 | openrouter | 1.05M | $3 | $15 |
minimax/minimax-m3 | openrouter | 1.05M | $0.3 | $1.2 |
deepseek/deepseek-v4-pro | openrouter | 1.05M | $0.687648 | $1.375296 |
upstage/solar-pro4 | openrouter | 524K | $0.03 | $0.12 |
deepseek/deepseek-v4-pro-0813 | openrouter | 1.05M | $1.12068 | $3.36204 |
google/gemini-3.8-flash | openrouter | 1.05M | $0.75 | $3.75 |
A model or mix you deploy from Optimize gets an id of its own, tm/<name>-v<n>. It is callable on POST /v1/chat/completions with a key of the organization that owns it, and it is not listed in the catalog.
Image models
Image models bill in tokens like everything else, but they answer on POST /v1/images/generations and nowhere else. The per-image ceiling is the most output tokens one image can bill; the reservation holds n times it.
| Model id | Sizes | Qualities | Per-image ceiling | Text in /1M | Image out /1M |
|---|---|---|---|---|---|
gpt-image-1 | 1024x1024 · 1536x1024 · 1024x1536 | low · medium · high | 6,240 | $5 | $40 |
gpt-image-1-mini | 1024x1024 · 1536x1024 · 1024x1536 | low · medium · high | 6,500 | $2 | $8 |
gpt-image-1.5 | 1024x1024 · 1536x1024 · 1024x1536 | low · medium · high | 6,500 | $5 | $32 |
gpt-image-2 | 1024x1024 · 1536x1024 · 1024x1536 | low · medium · high | 8,232 | $5 | $30 |
gpt-image-2.5-flare | 1024x1024 · 1536x1024 · 1024x1536 | low · medium · high · xhigh · max | 8,232 | $5 | $30 |
gpt-image-2.5-sunburst | 1024x1024 · 1536x1024 · 1024x1536 | low · medium · high · xhigh · max | 8,232 | $5 | $30 |
size and quality also accept auto, which lets the provider choose. Every model's full descriptor map is in /api/models.json under supported_parameters.
The Nano Banana, Seedream and Grok Imagine models are served through OpenRouter and billed at the cost it reports (how). They take a shape as aspect_ratio (16:9, 9:16, …) instead of a size, most take a resolution (1K, 2K, 4K), and Seedream takes a seed.
Video models
Video models answer on POST /v1/videos and nowhere else. A video is a job: it is created at once and made in the next seconds to minutes. All three are served through OpenRouter today and billed at the cost it reports.
| Model id | Made by | Length | Resolutions | Listed price | Per second ≤ |
|---|---|---|---|---|---|
seedance-2.5 | ByteDance | 4–30 s | 480p · 720p | $10.70 / M video tokens | $0.30 |
hailuo-3-max | MiniMax | 5–15 s | 480p · 768p | $0.05–$0.08 per second | $0.08 |
wan-3.0 | Alibaba | 5–30 s | 480p · 720p · 1080p | $0.05–$0.20 per second | $0.20 |
Decision models
Decision models answer on POST /v1/decisions and nowhere else. They write no text: they return one typed answer per question about a state. Each model takes its own question format.
| Model id | Made by | Providers | Question kinds | Input /1M |
|---|---|---|---|---|
typesafe/jev-1.13 | TypeSafe | typesafe, then openrouter-decisions | choice, score, yes/no (System One) | $0.042 |
inception/mercury-decide | Inception | openrouter-decisions-free | choice, score, yes/no (System One, the same as Jev) | $0 |
bespokelabs/nimble-v3 | Bespoke Labs | bespokelabs | choice, score, yes/no (System One, the same as Jev) | $0.04 |
cloudflare/clef | Cloudflare | workers-ai | choice, score, yes/no (System One, the same as Jev) | $0.24 |
cloudflare/clef-flash | Cloudflare | workers-ai | choice, score, yes/no (System One, the same as Jev) | $0.09 |
routerplus/decider-2b | RouterPlus | routerplus | choice, score, yes/no (System One, the same as Jev) | $0 |
routerplus/kev-4b | RouterPlus | routerplus | choice, score, yes/no (System One, the same as Jev) | $0 |
perplexity/pplx-decider-v1-27b | Perplexity | perplexity-decisions | choice, score, yes/no (System One, the same as Jev) | $0.04 |
levanto/sage-1.2 | Levanto | levanto | choice, score, yes/no (System One, the same as Jev) | $0.05 |
fastino/gliner-2.5-decide | Fastino | fastino | classifications (single- and multi-label), entities, structures, relations | $0.03 |
Output is free on every decision model. Sage, like Perplexity Decider, bills the state once per question. Mercury Decide is $0 in and out while OpenRouter serves only its free variant; its pool takes 20 requests a minute (15 per organization) and 5 in flight at once (3 per organization), and OpenRouter caps free requests per day. Bespoke Nimble v3 is $0.04 per million input tokens, Bespoke's own price with no markup. It takes 1 to 64 questions per call, and its pool takes 7 requests at once (5 per organization): Bespoke allows our account 8, and one is kept for our test environment. Clef and Clef-flash are Cloudflare's own decision models, served by Cloudflare Workers AI at Cloudflare's own price with no markup: $0.24 and $0.09 per million input tokens. They take 1 to 64 questions per call, each question id 1 to 100 letters, digits, _, . or -, and a context of 65,536 tokens. They share one pool: 200 requests a minute (150 per organization) and at most 8 at a time (6 per organization). Decider 2B and Kev 4B are RouterPlus's own models, on our own GPUs in the United States, and free: $0 in and out. They share one pool. Their contexts are 25,600 and 8,192 tokens. RouterPlus answers an identical request from its cache for 600 s. Perplexity Decider v1 27B is Perplexity's own decision model, served by Perplexity's API at Perplexity's own price with no markup: $0.04 per million input tokens. Perplexity bills the state once per question, so a call costs about one state for each question it asks. It takes 1 to 128 questions per call, and 262,144 tokens for each question's prompt. Its pool takes 540 requests a minute (405 per organization) and 8 at once (6 per organization). A discount may apply to decision models; see Pricing.
Most models also carry cache pricing (cached_prompt, and cache_write where the provider bills it) — the full per-SKU table is on each model's page and in /api/models.json. A model with no separate cache or reasoning price bills those tokens at the base rate — a missing SKU never means free.
One model, several providers
A model id can be served by more than one provider — a first-party lab and an aggregator, say. The catalog shows one card; the model page lists every provider with its own uptime and TTFB. Routing tries providers in priority order and fails over before the first byte only — never mid-answer. Prices are identical across providers of one model (enforced when the catalog is built), so which provider served you never changes your bill. The x-tm-provider header says who did; when that provider knows the model under its own id (an aggregator's anthropic/claude-sonnet-4.5 for our claude-sonnet-4-5), x-tm-upstream-model carries the id that was sent and the response's model field echoes the provider's id truthfully. When an aggregator names the provider it used, x-tm-served-by carries that name (for example Amazon Bedrock).
Date-pinned ids
The catalog lists one clean id per model (claude-haiku-4-5), but vendors also ship dated snapshots (claude-haiku-4-5-20251001, gpt-4o-2024-08-06) and -latest spellings, and tools like Claude Code send them. Any such pin of a catalog model works without being listed: the gateway resolves it to its family for routing and pricing only, and forwards your original id to the provider verbatim — a snapshot pin is never silently retargeted to a newer version. The response's model field shows the exact snapshot that served you. If a vendor retires a pinned snapshot, you get the vendor's own error, not a quiet upgrade. Unknown ids still 404 naming themselves, dated or not. A key bound to your own provider connection matches ids literally, so list the exact ids you send there.
Id grammar
Ids match [a-zA-Z0-9][a-zA-Z0-9._/-]{0,127} — letters, digits, . _ / -, max 128 chars. : is deliberately excluded: the suffix namespace (:something) is reserved for future marketplace variants, so a provider can never squat a routing suffix.
Providers are manifests
A provider is a JSON manifest in the repo — endpoint, wire dialect, settlement terms, data policy, and a model list with exact prices:
{
"provider": { "id": "anthropic", "name": "Anthropic (direct API)", "prompt_logging": "retained" },
"endpoint": { "base_url": "https://api.anthropic.com/v1", "dialect": "anthropic" },
"models": [{
"id": "claude-haiku-4-5",
"context_length": 200000,
"max_output_tokens": 64000,
"pricing": [
{ "type": "prompt", "unit": "token", "cost_usd_per_million": "1" },
{ "type": "completion", "unit": "token", "cost_usd_per_million": "5" }
]
}]
}Manifests compile into the routing catalog through a gated pipeline — validate, emit a staging catalog, run the live conformance suite, then promote. Adding or repricing a model is a data change with a gate, never a gateway code deploy. Three lifecycle rules do real work:
is_ready: falsestages a model: validated and testable, but never routed and never listed.deprecation_datedelists automatically: past that date the model drops out of the compiled catalog and the public pages.- Shared ids must agree. If two providers declare the same model id with different prices, or one as a chat model and the other as an image model, the catalog build fails. Pass-through pricing stays coherent per id.
Prices sync as effective-dated append-only rows. A price change never rewrites history — old requests stay priced as dispatched.
Conformance before listing
No model is listed on a provider's say-so. The conformance suite runs live requests against the provider's endpoint, and passing it is the gate between staging and live. It is deliberately implemented independently of the gateway's own wire parsers, so a shared bug can't hide real breakage.
| Check | Proves |
|---|---|
| C1 | A non-stream completion returns real usage tokens (usage is the billable record) |
| C2 | Stream frames are parseable SSE |
| C3 | Streams terminate properly ([DONE] / message_stop) |
| C4 | Usage tokens are reported in-stream |
| C5 | An unknown model yields a parseable 4xx — not a silent fallback |
| C6 | Responses carry actual assistant content, not just billable counters |
| C7 | The response echoes the requested model id — no silent substitution |
| I1 | An image generation at the top declared quality returns base64 data and usage |
| I2 | The declared per-image ceiling covers the billed output |
| I3 | An undeclared size is refused |
| I4 | The ceiling holds at every declared size (release-day probe, opt-in: it buys a real render per size) |
| V1 | A video job at the cheapest tier finishes, reports its cost and downloads as an MP4 |
| V2 | The declared per-second ceiling covers the reported cost |
| V3 | An undeclared shape is refused |
| V4 | The ceiling holds at the top resolution (opt-in: it buys one more job per model) |
Image listings run I1–I3 and C5; video listings run V1–V3 and C5. C2–C4, C6 and C7 do not apply to them — the media routes do not stream, carry no assistant text, and an images response has no model field to echo.
C6 and C7 exist because the failure modes that matter are billing for an empty answer and quietly serving a cheaper model. C7 accepts exactly one alias form: a dated snapshot of the same id (gpt-4o → gpt-4o-2024-08-06), or an alias target the manifest declares up front in resolves_to. gpt-4o answered by gpt-4o-mini fails.
The directory
https://app.routerplus.com/ lists every live model in three sections — chat, image and video — because they are called on different routes. A row shows the model's context length and max output (for a chat model), or its sizes, qualities, lengths and resolutions (for an image or video model), plus its 24-hour request count. The list can be searched (/?q=haiku matches model, lab and provider names), filtered by kind, context length, supported parameters, zero data retention, input price, lab and provider, and sorted by use, name, price or context. Prices, latency and uptime depend on the provider, so they live on each model's page at /models/<id>: one row per provider that runs the model, per-SKU pricing, the provider's prompt-retention policy (marketplace logs are content-free either way), and ready-to-paste snippets.
The uptime figure is allowed to say nothing
Uptime and latency come from the trailing 24 hours of real routed traffic. When a provider we call directly has fewer than 100 counted attempts in that window, the model page shows n<100 instead of a percentage. For a provider reached through an aggregator, the figure is the one the aggregator publishes.
A model with 3 requests is not "100% up", and we won't render it that way. No number until the sample is real — that rule is load-bearing, not a placeholder.
The hourly bars on a model page are green at ≥ 99%, amber at ≥ 95%, and red below that; an hour with no sample stays grey.
Machine-readable: /api/models.json
Everything above, as JSON, no auth:
curl -s https://app.routerplus.com/api/models.json{
"models": [
{
"id": "claude-haiku-4-5",
"display_name": "Claude Haiku 4.5",
"output_modalities": ["text"],
"providers": [
{ "id": "anthropic", "name": "Anthropic (direct API)", "dialect": "anthropic", "prompt_logging": "retained" },
{ "id": "openrouter", "name": "OpenRouter", "dialect": "openai", "prompt_logging": "retained",
"upstream_id": "anthropic/claude-haiku-4.5" }
],
"served_by": [
{ "id": "anthropic", "name": "Anthropic", "zdr": false,
"routes": [ { "via": "anthropic", "direct": true }, { "via": "openrouter", "direct": false } ] },
{ "id": "amazon", "name": "Amazon", "zdr": true, "routes": [ { "via": "openrouter", "direct": false } ] },
{ "id": "azure", "name": "Azure", "zdr": false, "routes": [ { "via": "openrouter", "direct": false } ] },
{ "id": "google-vertex", "name": "Google Vertex", "zdr": true, "routes": [ { "via": "openrouter", "direct": false } ] }
],
"context_length": 200000,
"max_output_tokens": 64000,
"pricing_usd_per_million": {
"prompt": "1", "cached_prompt": "0.10",
"cache_write": "1.25", "completion": "5"
},
"lab": { "id": "anthropic", "name": "Anthropic" },
"weights": "proprietary"
}
],
"gateway": "https://api.routerplus.com/v1",
"docs": "https://app.routerplus.com/docs"
}| Field | Meaning |
|---|---|
id | The exact string to put in your request's model field |
output_modalities | ["text"], ["image"], ["video"] or ["decisions"] — which route serves this model. Filter on it; never guess from the id |
providers | Every deployment serving this id, primary first, each with id, name, dialect, prompt_logging, and upstream_id when the provider knows the model under its own name |
dialect (per provider) | The provider's native wire format — informational; every chat model answers on both surfaces |
prompt_logging (per provider) | Provider's prompt retention: none or retained |
served_by | The providers that run the model, one entry per provider however it is reached. routes[].via is the providers entry the request goes to; direct: false means an aggregator picks this provider per request. quantizations appears when the aggregator publishes it (for example ["fp8"]). zdr is true when the provider offers Zero Data Retention for this model |
context_length / max_output_tokens | Token limits; max_output_tokens may be null, and context_length is null for image and video models |
supported_parameters | Image and video models only: the descriptor map the route enforces (size, quality, n, seconds, and the rest) |
pricing_usd_per_million | Decimal-string USD per 1M tokens, keyed by SKU type |
billing | "reported_cost" on the models billed at the provider's reported cost; absent on token-billed models |
price_card | On those models, the provider's own listed prices, for reading — see Pricing & billing |
lab | Who made the model, as distinct from who serves it — null when unknown |
weights | "open" when the lab publishes the model's weights, "proprietary" when it never has, null when we are not certain |
For an image model max_output_tokens is the per-image ceiling, not a cap on a completion. For a video model it is the ceiling per second of video, in cost units (one micro-dollar each).
Prices are decimal strings so nothing is lost to float rounding — parse them with a decimal type if you're doing money math. For the authenticated, SDK-shaped list (what client.models.list() calls), see GET /v1/models.
Markdown source for agents: /docs/models.md · index at /llms.txt