# Pricing & billing

Pricing is pass-through: the per-token price you pay is the provider's price,
published per model in USD per million tokens. There is no per-token markup and
no hidden fee line. Every billed response carries `usage.cost` — the exact
amount your balance was debited, computed with the same integer math the ledger
settles with — so you can recompute your bill from the wire at any time. One provider
is our own: RouterPlus, which serves our own decision models on our own GPUs. Its models are
free: $0 in and out (see [Decision models](#decision-models)).

## Where prices live

The machine-readable catalog is public, no key required:

```bash
curl -s https://app.routerplus.com/api/models.json \
  | jq '.models[] | select(.id == "claude-sonnet-4-5").pricing_usd_per_million'
```

```json
{
  "prompt": "3",
  "cached_prompt": "0.30",
  "cache_write": "3.75",
  "completion": "15"
}
```

Prices are **decimal strings** (up to six fractional digits), USD per one
million tokens. They convert exactly to the ledger's internal unit — no float
rounding is involved anywhere in billing.

| SKU | What it bills | Example (`claude-sonnet-4-5`) |
|---|---|---|
| `prompt` | non-cached input tokens | $3 / M |
| `cached_prompt` | prompt-cache **reads** | $0.30 / M |
| `cache_write` | prompt-cache **writes** | $3.75 / M |
| `completion` | output tokens (including reasoning, unless a distinct rate is declared) | $15 / M |
| `internal_reasoning` | reasoning subset of output, when a provider declares a distinct rate | — |

> [!NOTE]
> A model that does not declare a cache or reasoning SKU bills those tokens at
> its base rate: `cached_prompt` and `cache_write` default to the `prompt`
> price, `internal_reasoning` defaults to the `completion` price. A missing
> price never means free — that would be an accidental $0 SKU, and the ledger
> refuses to guess.

## Image models

Image models bill on the **same two SKUs**: `prompt` for the text you send in,
`completion` for the image that comes out, both in USD per million tokens.
OpenAI reports image output in tokens, and the count is fixed by the size and
quality you ask for. The per-image ceiling — the most output tokens one image
can bill — is `max_output_tokens` on the model page and in
`/api/models.json`.

| Model id | Text in /1M | Image out /1M | Per-image ceiling |
|---|---|---|---|
| `gpt-image-1` | $5 | $40 | 6,240 tokens |
| `gpt-image-1-mini` | $2 | $8 | 6,500 tokens |
| `gpt-image-1.5` | $5 | $32 | 6,500 tokens |
| `gpt-image-2` | $5 | $30 | 8,232 tokens |
| `gpt-image-2.5-flare` | $5 | $30 | 8,232 tokens |
| `gpt-image-2.5-sunburst` | $5 | $30 | 8,232 tokens |

Every output token bills at the image rate, including the text output tokens
`gpt-image-1.5` folds into its `output_tokens`. The full route contract is
[POST /v1/images/generations](/docs/api-images).

## Models billed at the provider's reported cost

Some image and video models are served through OpenRouter today: Google's Nano Banana
family, ByteDance's Seedream and Seedance, xAI's Grok Imagine, MiniMax's H3 Max and
Alibaba's Wan 3.0. Their providers price them per image, per second of video or per video
token — not in text tokens. For these models **the charge is exactly the cost the provider
reports** for the request, passed through with no margin.

- The models page and `/api/models.json` show each model's own listed prices
  (`price_card`).
- The reservation is a ceiling: per image for an image model, per second for a video model.
  The models page shows it as **Per image ≤** or **Per second ≤**.
- When the provider reports no cost, the charge is the ceiling, marked `estimated`.
- A reported cost above the ceiling is still charged in full: the ceiling limits what is
  held, not what is billed.

In the ledger these rows count *cost units*: one unit is one micro-dollar, priced at $0 per
million in and $1 per million out. So `output_tokens` on such a row is the charge in
micro-dollars, and the per-million prices in `/api/models.json` read `"0"` and `"1"` for
these models. The response you read still carries the provider's own token counts, and
`usage.cost` in USD.

| Model id | Kind | Listed price | Ceiling |
|---|---|---|---|
| `gemini-2.5-flash-image` | image | $30 / M image tokens | $0.05 per image |
| `gemini-3.1-flash-image` | image | $60 / M image tokens | $0.20 per image |
| `gemini-3-pro-image` | image | $120 / M image tokens | $0.30 per image |
| `seedream-5.0-pro` | image | $0.045 per image; $0.09 at 2K | $0.10 per image |
| `seedream-5.0-lite` | image | $0.035 per image | $0.04 per image |
| `grok-imagine-image-2.0` | image | $0.04–$0.08 per image, by quality and resolution | $0.09 per image |
| `seedance-2.5` | video | $10.70 / M video tokens | $0.30 per second |
| `hailuo-3-max` | video | $0.05 per second at 480p, $0.08 at 768p | $0.08 per second |
| `wan-3.0` | video | $0.05, $0.10, $0.20 per second at 480p, 720p, 1080p | $0.20 per second |

The route contracts are [POST /v1/images/generations](/docs/api-images) and
[POST /v1/videos](/docs/api-videos).

## Decision models

Decision models answer typed questions on [POST /v1/decisions](/docs/api-decisions). They
are priced like everything else: USD per million tokens.

| Model id | Input /1M | Output /1M | What the token counts are |
|---|---|---|---|
| `typesafe/jev-1.13` | $0.042 | $0 | The provider's input count |
| `inception/mercury-decide` | $0 | $0 | The provider's input count |
| `bespokelabs/nimble-v3` | $0.04 | $0 | The provider's input count |
| `cloudflare/clef` | $0.24 | $0 | The provider's input count |
| `cloudflare/clef-flash` | $0.09 | $0 | The provider's input count |
| `routerplus/decider-2b` | $0 | $0 | The provider's input count; an answer from its cache bills $0 |
| `routerplus/kev-4b` | $0 | $0 | The provider's input count; an answer from its cache bills $0 |
| `perplexity/pplx-decider-v1-27b` | $0.04 | $0 | The provider's input count, which holds the state once per question |
| `levanto/sage-1.2` | $0.05 | $0 | The provider's input count, which holds the state once per question |
| `fastino/gliner-2.5-decide` | $0.03 | $0 | Fastino's `prompt_tokens` |

Mercury Decide is $0 in and $0 out because OpenRouter, its only provider, serves only the
model's free variant today, and pricing is pass-through. `usage.cost` is `0` on every call.
A paid variant, or Inception serving the model directly, would be a new listed price, shown
on the model's page and in `/api/models.json` like any other. A free call is still metered,
counted against your limits and its pool, and recorded in `GET /v1/generation`, and it still
needs an organization with credit (see the discount note below).

Bespoke Nimble v3 is billed on Bespoke's own input count at Bespoke's own price, with no
markup: $0.04 per million input tokens, and output is free. Bespoke counts the state and
the questions once per call, not once per question. A three-question call of about 340
input tokens costs $0.000013.

Clef and Clef-flash are billed on Cloudflare's own input count at Cloudflare's own price,
with no markup: $0.24 and $0.09 per million input tokens, and output is free (Cloudflare
reports `output_tokens: 0`). Cloudflare bills Workers AI in neurons, at $0.011 per 1,000
neurons. Its price page lists Clef at 21,818 neurons ($0.24) and Clef-flash at 8,182 neurons
($0.09) per million input tokens. A call of 400 input tokens costs $0.000096 on Clef and
$0.000036 on Clef-flash.

Decider 2B and Kev 4B are RouterPlus's own models, on our own GPUs, not an outside
provider's. So their price is ours to set, not a provider's price passed through, and it is
$0: input, cached input and output are all free, and `usage.cost` is `0` on every call.
RouterPlus still reports its input count, which is at most the model's context: a long input
that Decider 2B (25,600 tokens) cuts counts only the tokens the model read. A free call is
still metered, counted against your limits and its pool, and recorded in
`GET /v1/generation`, and it still needs an organization with credit. RouterPlus answers an
identical request from its cache for 600 s.

Perplexity Decider v1 27B is billed on Perplexity's own input count at Perplexity's own
price, with no markup: $0.04 per million input tokens, and output is free. Unlike Bespoke,
Perplexity runs one prompt for each question, each with the whole state, and counts the
state once per question. So a call costs about one state for each question it asks: five
questions about a state of 1,843 tokens bill 9,215 input tokens, $0.000368. One call with
many questions still saves round trips.

GLiNER-2.5-Decide is billed on Fastino's own input count: $0.03 per million input tokens,
and output is free. `usage.input_tokens` is Fastino's `prompt_tokens`, and
`usage.output_tokens` its `completion_tokens`. A call of 1,000 input tokens costs $0.00003.

Sage is billed on Levanto's own input count at Levanto's own price, with no markup: $0.05
per million input tokens, and output is free. Like Perplexity Decider, it reads and bills the
state once per question, so a call costs about one state for each question it asks.

### The decisions discount

A discount may apply to decision models: a percent taken off the charge of every decision
model, or of some of them, each at its own percent, from every key, the playground's
included. It does not change the list price, and it can change or end. A discounted model's
page and its row in the directory show the percent beside the struck list price, and its row
in `/api/models.json` carries `discount_percent`. Every discounted response says what
applied:

| Where | What it shows |
|---|---|
| `usage.cost` | What you paid, after the discount |
| `usage.cost_before_discount` | The charge at the list price |
| `usage.discount_percent` and the `x-tm-discount-percent` header | The percent taken off |
| `cost_usd` in [GET /v1/generation](/docs/api-usage) | What you paid, after the discount |
| `inference_cost_usd` in GET /v1/generation | The cost at the list price |

The discounted charge is the list charge × (100 − percent) / 100, rounded down to the
micro-dollar. When no discount applies, the three discount fields are absent and
`usage.cost` is the list charge. A discount is a lower charge, not a markup: pricing stays
pass-through. A discount is for organizations with credit: with a balance of $0 a
discounted request is refused with `429 insufficient_quota`, as any request is, even at
100 %.

When a discount covers a model whose list price is already $0 (Mercury Decide, Decider 2B and Kev 4B today), the discount has nothing to
take off: `usage.cost` and `usage.cost_before_discount` are both `0`. The gateway still
reports `usage.discount_percent` and the header on it, as on every model the discount
covers, and `/api/models.json` still carries `discount_percent` on its row, so one reader works for all
ten models. The model's page, the directory and the Playground show no "N % off" beside a
$0 price: the model is free, not discounted.

## The math: integer micro-USD

All money math is integer micro-USD (millionths of a dollar). A published price
like `"2.50"` becomes exactly `2500000` micro-USD per million tokens. Two
rounding rules, both fixed:

- **Reserve rounds up** (ceiling) — the pre-flight hold is conservative.
- **Settlement rounds down** (floor) — the actual charge, in your favor.

The settlement formula, verbatim from the ledger:

```
non_cached_input = max(0, input_tokens − cached_tokens − cache_write_tokens)
reasoning        = clamp(reasoning_tokens, 0, output_tokens)
plain_output     = output_tokens − reasoning

cost_micro = floor((
    non_cached_input   × prompt_per_M
  + cached_tokens      × cached_prompt_per_M
  + cache_write_tokens × cache_write_per_M
  + plain_output       × completion_per_M
  + reasoning          × reasoning_per_M
) / 1,000,000)
```

where each `*_per_M` is the integer micro-USD price. `usage.cost` on the wire
is this number divided by 10⁶.

## Reserve, then settle

Every billed request runs the same lifecycle:

1. **Reserve.** Before any provider is contacted, the gateway holds a
   conservative worst case. On the chat surfaces the input is estimated at one
   token per byte of the request body, plus 65,536 tokens for each image or
   document part (an image sent inline counts the 65,536 only, not its base64
   bytes), and the output at your `max_tokens` (or
   `max_completion_tokens`; 4096 when omitted, and never more than 32,768 —
   see [POST /v1/chat/completions](/docs/api-chat-completions)). Each side is
   priced at the model's most expensive rate for that side — cache writes and
   reasoning when they cost more than plain input or output — and the sum is
   rounded **up** to micro-USD. On `/v1/images/generations` the output is
   reserved at `n` × the model's per-image ceiling (not `max_tokens`), and the
   input at the prompt bytes ÷ 4; on `/v1/videos` the hold is the length in
   seconds × the model's per-second ceiling.
2. **Dispatch.** The request goes to a provider. Prices are pinned at this
   moment (see [frozen snapshots](#price-changes-frozen-snapshots)).
3. **Settle.** When the attempt finishes, the provider-reported token counts
   are priced with the formula above, rounded **down**. That is what you pay.
4. **Release.** The hold is released in full. You are never charged the
   reserved maximum — with one honest exception: if the gateway crashes
   mid-request and cannot observe the outcome, the attempt settles at the
   reservation, because the call may have been billed upstream. That outcome is
   visible as `unknown_after_crash` in
   [GET /v1/usage & /v1/generation](/docs/api-usage).

A reservation can be refused before dispatch:

- **Balance too low** — your balance must cover every reservation your org has
  in flight, including this one. Refusal is a `429` with
  `error_type: insufficient_quota` (the OpenAI-native shape SDKs already
  handle; `billing_error` native type on the Anthropic surface).
- **Monthly spend cap hit** — hard caps (org-wide and per-key, UTC calendar
  month, strictest wins) return the same `429 insufficient_quota` with the
  exact reset time in the message and the `x-tm-cap-reset` header.
- **A rate limit hit** — the key's RPM, a shared limit of the organization, or
  the provider's pool: `429` with `error_type: rate_limit` and `retry-after`.

Remediation for each is in [Errors & remediation](/docs/errors).

### All-or-nothing on images

An image generation either arrives whole or costs nothing. A failed attempt, a
response over the 32 MiB cap, and a 200 that carries no decodable image all
settle at zero — there is no partial image and no partial bill. The one nuance:
if you hang up mid-render the render still finishes upstream, so that attempt
settles the **provider-reported** usage, never the reservation and never a
pretend $0.

## The cost on the wire: `usage.cost`

Every billed response carries `usage.cost` (USD, a JSON number):

- **Non-streaming** — inline in the response body's `usage`, both surfaces.
- **Streaming, OpenAI surface** — in the final usage chunk, before `[DONE]`.
- **Streaming, Anthropic surface** — in the `message_delta` usage at stream end.

```bash
curl -s https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-4-5","max_tokens":350,"messages":[{"role":"user","content":"hello"}]}' \
  | jq .usage
```

```json
{
  "prompt_tokens": 1200,
  "completion_tokens": 350,
  "total_tokens": 1550,
  "prompt_tokens_details": { "cached_tokens": 800, "cache_write_tokens": 0 },
  "cost": 0.00669
}
```

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])
r = client.chat.completions.create(
    model="claude-sonnet-4-5", max_tokens=350,
    messages=[{"role": "user", "content": "hello"}],
)
print(r.usage.model_dump()["cost"])  # the exact ledger debit, in USD
```

```python
import os
import anthropic

client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])
msg = client.messages.create(
    model="claude-sonnet-4-5", max_tokens=350,
    messages=[{"role": "user", "content": "hello"}],
)
print(msg.usage.model_dump()["cost"])
```

> [!WARNING]
> `usage.cost` is the **full request debit**, not the serving attempt's slice.
> If a provider failed mid-prompt and the gateway failed over before your
> answer started, any tokens the failed attempt consumed are included in the
> number you see. The per-attempt breakdown is always available at
> [GET /v1/generation](/docs/api-usage) — nothing is hidden in an average.

One documented limitation: an Anthropic-surface **stream that dies after your
answer started** has no legal wire slot for usage outside `message_delta`, so
its terminal error event carries no cost. Recompute those from
[/v1/generation](/docs/api-usage). (The OpenAI surface emits any known
usage+cost before its terminal error, so its wire stays complete.)

Usage reporting is always on: streams and non-streams both carry it, and the
gateway requests usage upstream regardless of what you send, because
unreported usage would be an unrecomputable bill. The one `stream_options` value
that changes anything is an explicit `include_usage: false`: you are billed the
same, but no usage chunk is written to you.

## Cache reads and writes

Cache tokens are billed at their own rates on **both** surfaces, and both
surfaces report the split:

| Concept | OpenAI surface | Anthropic surface | Billed at |
|---|---|---|---|
| Non-cached input | `prompt_tokens` minus the two details below | `input_tokens` | `prompt` |
| Cache reads | `prompt_tokens_details.cached_tokens` | `cache_read_input_tokens` | `cached_prompt` |
| Cache writes | `prompt_tokens_details.cache_write_tokens` | `cache_creation_input_tokens` | `cache_write` |
| Output | `completion_tokens` | `output_tokens` | `completion` |

On the OpenAI surface `prompt_tokens` is cache-inclusive (reads and writes are
inside it), and billed responses always carry `prompt_tokens_details` with both
fields — zero-filled when the upstream reports none — so stream and non-stream
usage share one recomputable shape.

## Reasoning tokens

Reasoning tokens are a **subset of output**, never billed on top of it. The
settle math splits them out and prices them at the model's reasoning rate;
today every cataloged model prices reasoning equal to `completion`, so the
split is billing-neutral until a provider declares a distinct rate. The OpenAI
surface reports them in `completion_tokens_details.reasoning_tokens`; the
Anthropic wire does not report them separately.

## Price changes: frozen snapshots

The price you pay is the price at **dispatch time**. Price rows are
append-only — a change is a new effective-dated row, never an edit — and the
row in effect when your request is reserved is pinned into the reservation and
written onto every attempt's ledger row. A price change published while your
request is in flight cannot touch it, and historical attempts always settle
(and audit) against the prices they were dispatched under.

## Paid credits and purchase matching

Create an account at [Sign up](https://app.routerplus.com/signup), complete Clerk authentication
and email verification, and save your first API key. Signup and card verification
alone do not grant credit. Existing customers keep their balance and API keys.

Standard credit purchases start at **$10** in [Billing](https://app.routerplus.com/console/billing).
The payment adds the amount you buy as paid credit. Your balance is the sum of
credit ledger entries minus settled spend.

We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match.

With the full match available: $10 purchase → $20 credits. $100 purchase → $200 credits. A partially used cap adds only the matching credit remaining.

The cap applies across eligible purchases, not afresh to each checkout. Billing
shows the offer and remaining matching credit for your account. Matching credit
is promotional and appears separately in the ledger. Refunds and chargebacks
reverse the corresponding match; a partial refund reverses its proportional
matching credit. Adding a card without a
purchase does not earn a bonus. Existing trial and card-verification grants stay
in the ledger, but those signup offers are no longer available.

Before an organization buys credit, its customer keys share the trial rate of
20 requests/minute. A purchase lifts that cap; each key then has its own ceiling
(300 for a console key). Historical trial credit can also be limited when unusual
usage is detected; contact support if that happens.

## Worked example

A `claude-sonnet-4-5` request reports 1,200 prompt tokens (800 of them cache
reads, no cache writes) and 350 completion tokens:

```
non-cached input:   400 × 3,000,000  = 1,200,000,000
cache reads:        800 ×   300,000  =   240,000,000
output:             350 × 15,000,000 = 5,250,000,000
                                       ─────────────
sum / 1,000,000 (floor)              = 6,690 micro-USD
```

`usage.cost` on the wire: `0.00669`. The same number appears as the attempt's
`cost_usd` in [GET /v1/usage & /v1/generation](/docs/api-usage), because they
are the same computation on the same integers.

## Worked example: an image

A `gpt-image-1` request for two medium 1024×1024 images (`n: 2`) with a
12-token prompt estimate. Text in is $5 / M, image out is $40 / M, and the
per-image ceiling is 6,240 tokens.

The reservation, before any provider is contacted:

```
prompt estimate:  12 ×  5,000,000 =         60,000,000
output ceiling:   2 × 6,240 × 40,000,000 = 499,200,000,000
                                           ───────────────
ceil(sum / 1,000,000)                    = 499,260 micro-USD
```

The provider reports 12 input tokens and 2,112 output tokens (1,056 per image),
so settlement is:

```
input:            12 ×  5,000,000 =         60,000,000
output:        2,112 × 40,000,000 =     84,480,000,000
                                        ──────────────
floor(sum / 1,000,000)            =     84,540 micro-USD
```

`usage.cost` on the wire: `0.08454`. The hold is released in full; you pay the
settled number, not the reservation.
