Console
Core concepts/Models & catalog

Models & catalog

The catalog, model ids, uptime honesty, machine-readable feeds.

/llms.txt

The catalog is one flat namespace of model ids. Every chat model in it is callable through both wire surfaces — OpenAI-compatible /v1/chat/completions and Anthropic-compatible /v1/messages — with one key; the gateway translates. Image models answer on POST /v1/images/generations only, video models on POST /v1/videos only, and decision models on POST /v1/decisions only. Pricing is pass-through: the per-token rates in the catalog are what you pay, and every billed response carries usage.cost so you can recompute the bill from the wire.

There are four ways to read the catalog:

SurfaceURLAuthFor
Model directoryhttps://app.routerplus.com/noneBrowsing, filters, 24-hour request counts
Per-model pageshttps://app.routerplus.com/models/<id>noneProviders with prices, latency, uptime and data policy; copy-paste snippets
Machine-readable JSONhttps://app.routerplus.com/api/models.jsonnoneScripts, dashboards, price checks
SDK-shaped listhttps://api.routerplus.com/v1/modelsAPI keyclient.models.list() — see GET /v1/models

Model ids

Ids are exact strings. There is no fuzzy matching and no silent substitution: a request for an id that is not in the catalog is a 404 that echoes the id you sent (see GET /v1/models). The catalog as of this writing — the live list is always /api/models.json. Models served by two providers list both, first-party first — see One model, several providers:

Model idProvidersContextInput /1MOutput /1M
claude-sonnet-5-5anthropic, openrouter1M$2$10
claude-opus-5-5anthropic, openrouter1M$4$20
claude-fable-5-1anthropic, openrouter1M$10$50
gpt-6.1-solopenrouter1.05M$2$10
gpt-6-solopenrouter1.05M$2$10
gpt-6-lunaopenrouter1.05M$0.1$0.5
gpt-6-astraopenrouter1.05M$10$50
claude-sonnet-5anthropic, openrouter1M$2$10
claude-opus-5anthropic, openrouter1M$5$25
claude-fable-5anthropic, openrouter1M$10$50
claude-sonnet-4-5anthropic, openrouter200K$3$15
claude-haiku-4-5anthropic, openrouter200K$1$5
gpt-4o-miniopenai, openrouter128K$0.15$0.6
gpt-4.1-miniopenai, openrouter1.05M$0.4$1.6
gpt-4oopenai, openrouter128K$2.5$10
tencent/hy4-previewopenrouter1.05M$0.834$2.501
z-ai/glm-5.3-flashopenrouter1.31M$0.075$0.25
deepseek/deepseek-v4.1-flashopenrouter1.05M$0.3$1.2
deepseek/deepseek-v4-flash-0731openrouter1.31M$0.04998$0.09996
deepseek/deepseek-v4-flashopenrouter1.05M$0.08078$0.16156
tencent/hy3openrouter262K$0.132$0.528
z-ai/glm-5.3openrouter1.31M$1.4$4.4
xiaomi/mimo-v2.5openrouter1.05M$0.14$0.28
z-ai/glm-5.2openrouter1.05M$0.966$3.036
google/gemini-3.7-flashopenrouter1.05M$0.75$3.75
moonshotai/kimi-k3openrouter1.05M$3$15
minimax/minimax-m3openrouter1.05M$0.3$1.2
deepseek/deepseek-v4-proopenrouter1.05M$0.687648$1.375296
upstage/solar-pro4openrouter524K$0.03$0.12
deepseek/deepseek-v4-pro-0813openrouter1.05M$1.12068$3.36204
google/gemini-3.8-flashopenrouter1.05M$0.75$3.75

A model or mix you deploy from Optimize gets an id of its own, tm/<name>-v<n>. It is callable on POST /v1/chat/completions with a key of the organization that owns it, and it is not listed in the catalog.

Image models

Image models bill in tokens like everything else, but they answer on POST /v1/images/generations and nowhere else. The per-image ceiling is the most output tokens one image can bill; the reservation holds n times it.

Model idSizesQualitiesPer-image ceilingText in /1MImage out /1M
gpt-image-11024x1024 · 1536x1024 · 1024x1536low · medium · high6,240$5$40
gpt-image-1-mini1024x1024 · 1536x1024 · 1024x1536low · medium · high6,500$2$8
gpt-image-1.51024x1024 · 1536x1024 · 1024x1536low · medium · high6,500$5$32
gpt-image-21024x1024 · 1536x1024 · 1024x1536low · medium · high8,232$5$30
gpt-image-2.5-flare1024x1024 · 1536x1024 · 1024x1536low · medium · high · xhigh · max8,232$5$30
gpt-image-2.5-sunburst1024x1024 · 1536x1024 · 1024x1536low · medium · high · xhigh · max8,232$5$30

size and quality also accept auto, which lets the provider choose. Every model's full descriptor map is in /api/models.json under supported_parameters.

The Nano Banana, Seedream and Grok Imagine models are served through OpenRouter and billed at the cost it reports (how). They take a shape as aspect_ratio (16:9, 9:16, …) instead of a size, most take a resolution (1K, 2K, 4K), and Seedream takes a seed.

Video models

Video models answer on POST /v1/videos and nowhere else. A video is a job: it is created at once and made in the next seconds to minutes. All three are served through OpenRouter today and billed at the cost it reports.

Model idMade byLengthResolutionsListed pricePer second ≤
seedance-2.5ByteDance4–30 s480p · 720p$10.70 / M video tokens$0.30
hailuo-3-maxMiniMax5–15 s480p · 768p$0.05–$0.08 per second$0.08
wan-3.0Alibaba5–30 s480p · 720p · 1080p$0.05–$0.20 per second$0.20

Decision models

Decision models answer on POST /v1/decisions and nowhere else. They write no text: they return one typed answer per question about a state. Each model takes its own question format.

Model idMade byProvidersQuestion kindsInput /1M
typesafe/jev-1.13TypeSafetypesafe, then openrouter-decisionschoice, score, yes/no (System One)$0.042
inception/mercury-decideInceptionopenrouter-decisions-freechoice, score, yes/no (System One, the same as Jev)$0
bespokelabs/nimble-v3Bespoke Labsbespokelabschoice, score, yes/no (System One, the same as Jev)$0.04
cloudflare/clefCloudflareworkers-aichoice, score, yes/no (System One, the same as Jev)$0.24
cloudflare/clef-flashCloudflareworkers-aichoice, score, yes/no (System One, the same as Jev)$0.09
routerplus/decider-2bRouterPlusrouterpluschoice, score, yes/no (System One, the same as Jev)$0
routerplus/kev-4bRouterPlusrouterpluschoice, score, yes/no (System One, the same as Jev)$0
perplexity/pplx-decider-v1-27bPerplexityperplexity-decisionschoice, score, yes/no (System One, the same as Jev)$0.04
levanto/sage-1.2Levantolevantochoice, score, yes/no (System One, the same as Jev)$0.05
fastino/gliner-2.5-decideFastinofastinoclassifications (single- and multi-label), entities, structures, relations$0.03

Output is free on every decision model. Sage, like Perplexity Decider, bills the state once per question. Mercury Decide is $0 in and out while OpenRouter serves only its free variant; its pool takes 20 requests a minute (15 per organization) and 5 in flight at once (3 per organization), and OpenRouter caps free requests per day. Bespoke Nimble v3 is $0.04 per million input tokens, Bespoke's own price with no markup. It takes 1 to 64 questions per call, and its pool takes 7 requests at once (5 per organization): Bespoke allows our account 8, and one is kept for our test environment. Clef and Clef-flash are Cloudflare's own decision models, served by Cloudflare Workers AI at Cloudflare's own price with no markup: $0.24 and $0.09 per million input tokens. They take 1 to 64 questions per call, each question id 1 to 100 letters, digits, _, . or -, and a context of 65,536 tokens. They share one pool: 200 requests a minute (150 per organization) and at most 8 at a time (6 per organization). Decider 2B and Kev 4B are RouterPlus's own models, on our own GPUs in the United States, and free: $0 in and out. They share one pool. Their contexts are 25,600 and 8,192 tokens. RouterPlus answers an identical request from its cache for 600 s. Perplexity Decider v1 27B is Perplexity's own decision model, served by Perplexity's API at Perplexity's own price with no markup: $0.04 per million input tokens. Perplexity bills the state once per question, so a call costs about one state for each question it asks. It takes 1 to 128 questions per call, and 262,144 tokens for each question's prompt. Its pool takes 540 requests a minute (405 per organization) and 8 at once (6 per organization). A discount may apply to decision models; see Pricing.

Most models also carry cache pricing (cached_prompt, and cache_write where the provider bills it) — the full per-SKU table is on each model's page and in /api/models.json. A model with no separate cache or reasoning price bills those tokens at the base rate — a missing SKU never means free.

One model, several providers

A model id can be served by more than one provider — a first-party lab and an aggregator, say. The catalog shows one card; the model page lists every provider with its own uptime and TTFB. Routing tries providers in priority order and fails over before the first byte only — never mid-answer. Prices are identical across providers of one model (enforced when the catalog is built), so which provider served you never changes your bill. The x-tm-provider header says who did; when that provider knows the model under its own id (an aggregator's anthropic/claude-sonnet-4.5 for our claude-sonnet-4-5), x-tm-upstream-model carries the id that was sent and the response's model field echoes the provider's id truthfully. When an aggregator names the provider it used, x-tm-served-by carries that name (for example Amazon Bedrock).

Date-pinned ids

The catalog lists one clean id per model (claude-haiku-4-5), but vendors also ship dated snapshots (claude-haiku-4-5-20251001, gpt-4o-2024-08-06) and -latest spellings, and tools like Claude Code send them. Any such pin of a catalog model works without being listed: the gateway resolves it to its family for routing and pricing only, and forwards your original id to the provider verbatim — a snapshot pin is never silently retargeted to a newer version. The response's model field shows the exact snapshot that served you. If a vendor retires a pinned snapshot, you get the vendor's own error, not a quiet upgrade. Unknown ids still 404 naming themselves, dated or not. A key bound to your own provider connection matches ids literally, so list the exact ids you send there.

Id grammar

Ids match [a-zA-Z0-9][a-zA-Z0-9._/-]{0,127} — letters, digits, . _ / -, max 128 chars. : is deliberately excluded: the suffix namespace (:something) is reserved for future marketplace variants, so a provider can never squat a routing suffix.

Providers are manifests

A provider is a JSON manifest in the repo — endpoint, wire dialect, settlement terms, data policy, and a model list with exact prices:

json
{
  "provider": { "id": "anthropic", "name": "Anthropic (direct API)", "prompt_logging": "retained" },
  "endpoint": { "base_url": "https://api.anthropic.com/v1", "dialect": "anthropic" },
  "models": [{
    "id": "claude-haiku-4-5",
    "context_length": 200000,
    "max_output_tokens": 64000,
    "pricing": [
      { "type": "prompt", "unit": "token", "cost_usd_per_million": "1" },
      { "type": "completion", "unit": "token", "cost_usd_per_million": "5" }
    ]
  }]
}

Manifests compile into the routing catalog through a gated pipeline — validate, emit a staging catalog, run the live conformance suite, then promote. Adding or repricing a model is a data change with a gate, never a gateway code deploy. Three lifecycle rules do real work:

  • is_ready: false stages a model: validated and testable, but never routed and never listed.
  • deprecation_date delists automatically: past that date the model drops out of the compiled catalog and the public pages.
  • Shared ids must agree. If two providers declare the same model id with different prices, or one as a chat model and the other as an image model, the catalog build fails. Pass-through pricing stays coherent per id.
Note

Prices sync as effective-dated append-only rows. A price change never rewrites history — old requests stay priced as dispatched.

Conformance before listing

No model is listed on a provider's say-so. The conformance suite runs live requests against the provider's endpoint, and passing it is the gate between staging and live. It is deliberately implemented independently of the gateway's own wire parsers, so a shared bug can't hide real breakage.

CheckProves
C1A non-stream completion returns real usage tokens (usage is the billable record)
C2Stream frames are parseable SSE
C3Streams terminate properly ([DONE] / message_stop)
C4Usage tokens are reported in-stream
C5An unknown model yields a parseable 4xx — not a silent fallback
C6Responses carry actual assistant content, not just billable counters
C7The response echoes the requested model id — no silent substitution
I1An image generation at the top declared quality returns base64 data and usage
I2The declared per-image ceiling covers the billed output
I3An undeclared size is refused
I4The ceiling holds at every declared size (release-day probe, opt-in: it buys a real render per size)
V1A video job at the cheapest tier finishes, reports its cost and downloads as an MP4
V2The declared per-second ceiling covers the reported cost
V3An undeclared shape is refused
V4The ceiling holds at the top resolution (opt-in: it buys one more job per model)

Image listings run I1–I3 and C5; video listings run V1–V3 and C5. C2–C4, C6 and C7 do not apply to them — the media routes do not stream, carry no assistant text, and an images response has no model field to echo.

C6 and C7 exist because the failure modes that matter are billing for an empty answer and quietly serving a cheaper model. C7 accepts exactly one alias form: a dated snapshot of the same id (gpt-4o → gpt-4o-2024-08-06), or an alias target the manifest declares up front in resolves_to. gpt-4o answered by gpt-4o-mini fails.

The directory

https://app.routerplus.com/ lists every live model in three sections — chat, image and video — because they are called on different routes. A row shows the model's context length and max output (for a chat model), or its sizes, qualities, lengths and resolutions (for an image or video model), plus its 24-hour request count. The list can be searched (/?q=haiku matches model, lab and provider names), filtered by kind, context length, supported parameters, zero data retention, input price, lab and provider, and sorted by use, name, price or context. Prices, latency and uptime depend on the provider, so they live on each model's page at /models/<id>: one row per provider that runs the model, per-SKU pricing, the provider's prompt-retention policy (marketplace logs are content-free either way), and ready-to-paste snippets.

The uptime figure is allowed to say nothing

Uptime and latency come from the trailing 24 hours of real routed traffic. When a provider we call directly has fewer than 100 counted attempts in that window, the model page shows n<100 instead of a percentage. For a provider reached through an aggregator, the figure is the one the aggregator publishes.

Note

A model with 3 requests is not "100% up", and we won't render it that way. No number until the sample is real — that rule is load-bearing, not a placeholder.

The hourly bars on a model page are green at ≥ 99%, amber at ≥ 95%, and red below that; an hour with no sample stays grey.

Machine-readable: /api/models.json

Everything above, as JSON, no auth:

bash
curl -s https://app.routerplus.com/api/models.json
json
{
  "models": [
    {
      "id": "claude-haiku-4-5",
      "display_name": "Claude Haiku 4.5",
      "output_modalities": ["text"],
      "providers": [
        { "id": "anthropic", "name": "Anthropic (direct API)", "dialect": "anthropic", "prompt_logging": "retained" },
        { "id": "openrouter", "name": "OpenRouter", "dialect": "openai", "prompt_logging": "retained",
          "upstream_id": "anthropic/claude-haiku-4.5" }
      ],
      "served_by": [
        { "id": "anthropic", "name": "Anthropic", "zdr": false,
          "routes": [ { "via": "anthropic", "direct": true }, { "via": "openrouter", "direct": false } ] },
        { "id": "amazon", "name": "Amazon", "zdr": true, "routes": [ { "via": "openrouter", "direct": false } ] },
        { "id": "azure", "name": "Azure", "zdr": false, "routes": [ { "via": "openrouter", "direct": false } ] },
        { "id": "google-vertex", "name": "Google Vertex", "zdr": true, "routes": [ { "via": "openrouter", "direct": false } ] }
      ],
      "context_length": 200000,
      "max_output_tokens": 64000,
      "pricing_usd_per_million": {
        "prompt": "1", "cached_prompt": "0.10",
        "cache_write": "1.25", "completion": "5"
      },
      "lab": { "id": "anthropic", "name": "Anthropic" },
      "weights": "proprietary"
    }
  ],
  "gateway": "https://api.routerplus.com/v1",
  "docs": "https://app.routerplus.com/docs"
}
FieldMeaning
idThe exact string to put in your request's model field
output_modalities["text"], ["image"], ["video"] or ["decisions"] — which route serves this model. Filter on it; never guess from the id
providersEvery deployment serving this id, primary first, each with id, name, dialect, prompt_logging, and upstream_id when the provider knows the model under its own name
dialect (per provider)The provider's native wire format — informational; every chat model answers on both surfaces
prompt_logging (per provider)Provider's prompt retention: none or retained
served_byThe providers that run the model, one entry per provider however it is reached. routes[].via is the providers entry the request goes to; direct: false means an aggregator picks this provider per request. quantizations appears when the aggregator publishes it (for example ["fp8"]). zdr is true when the provider offers Zero Data Retention for this model
context_length / max_output_tokensToken limits; max_output_tokens may be null, and context_length is null for image and video models
supported_parametersImage and video models only: the descriptor map the route enforces (size, quality, n, seconds, and the rest)
pricing_usd_per_millionDecimal-string USD per 1M tokens, keyed by SKU type
billing"reported_cost" on the models billed at the provider's reported cost; absent on token-billed models
price_cardOn those models, the provider's own listed prices, for reading — see Pricing & billing
labWho made the model, as distinct from who serves it — null when unknown
weights"open" when the lab publishes the model's weights, "proprietary" when it never has, null when we are not certain

For an image model max_output_tokens is the per-image ceiling, not a cap on a completion. For a video model it is the ceiling per second of video, in cost units (one micro-dollar each).

Prices are decimal strings so nothing is lost to float rounding — parse them with a decimal type if you're doing money math. For the authenticated, SDK-shaped list (what client.models.list() calls), see GET /v1/models.

Markdown source for agents: /docs/models.md · index at /llms.txt