# GET /v1/usage & /v1/generation

Two read-only endpoints on the gateway answer "what have I spent?" and "what
exactly did this request cost, attempt by attempt?". They read the same ledger
your balance is derived from — there is no separate reporting pipeline that
could disagree with your bill.

Both are **metadata only**: token counts, outcomes, and costs. Prompts and
completions are never stored, so they cannot be returned.

## Authentication

Both endpoints require a billed key (the first key shown after verified sign-in, or one created in the console)
and are scoped to that key's org, workspace and principal. Send it either way:

```bash
curl -s https://api.routerplus.com/v1/usage -H "Authorization: Bearer $TM_API_KEY"
curl -s https://api.routerplus.com/v1/usage -H "x-api-key: $TM_API_KEY"
```

> [!NOTE]
> A static development key (a gateway running without the money plane) has no
> ledger to read, so on these routes it gets the same undifferentiated 404 as
> any unknown path.

## GET /v1/usage

The 20 most recent attempts made with the keys of your workspace and principal,
newest first. The org balance fields are filled for a key in the organization's
default workspace and default principal — every first key shown after verified sign-in or made in the
console — and `null` for a key bound to an explicit workspace or principal
(see [Workspaces and principals](/docs/identity)).

```bash
curl -s https://api.routerplus.com/v1/usage -H "Authorization: Bearer $TM_API_KEY" | jq
```

```json
{
  "balance_usd": 0.98971,
  "credited_usd": 1.0,
  "spent_usd": 0.01029,
  "recent_attempts": [
    {
      "request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33",
      "deployment": "provider-b",
      "model": "claude-sonnet-4-5",
      "outcome": "completed",
      "usage_provenance": "observed",
      "input_tokens": 1200,
      "output_tokens": 350,
      "cost_usd": 0.00669,
      "billing_source": "house",
      "inference_cost_usd": 0.00669,
      "inference_price_source": "platform_list",
      "price_snapshot_id": "5d2c…",
      "fee_policy_version": "house-pass-through-v1",
      "at": "2026-09-04T09:12:33.104Z"
    },
    {
      "request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33",
      "deployment": "provider-a",
      "model": "claude-sonnet-4-5",
      "outcome": "failed",
      "usage_provenance": "observed",
      "input_tokens": 1200,
      "output_tokens": 0,
      "cost_usd": 0.0036,
      "billing_source": "house",
      "inference_cost_usd": 0.0036,
      "inference_price_source": "platform_list",
      "price_snapshot_id": "5d2c…",
      "fee_policy_version": "house-pass-through-v1",
      "at": "2026-09-04T09:12:31.512Z"
    }
  ]
}
```

| Field | Type | Meaning |
|---|---|---|
| `balance_usd` | number \| null | `credited_usd − spent_usd`, derived from ledger rows on read — never a cached counter |
| `credited_usd` | number \| null | sum of all credit grants to the org |
| `spent_usd` | number \| null | sum of every settled attempt |
| `recent_attempts[].request_id` | string | the request's `x-request-id`; feed it to `/v1/generation` |
| `recent_attempts[].deployment` | string | which deployment served the attempt — matches the `x-tm-provider` response header; `routerplus` on a closed [dedicated endpoint](/docs/dedicated-endpoints#closed-endpoints) and on any endpoint's dedicated capacity |
| `recent_attempts[].model` | string | the model id you requested; on a closed dedicated endpoint, the endpoint id |
| `recent_attempts[].outcome` | string | see the [outcome table](#attempt-outcomes) |
| `recent_attempts[].usage_provenance` | string \| null | `observed`, `estimated` or `unknown`; `null` before settlement |
| `recent_attempts[].input_tokens` / `output_tokens` | number | provider-reported counts (input is cache-inclusive). On a model billed at the provider's reported cost, `output_tokens` is the charge in micro-dollars |
| `recent_attempts[].cost_usd` | number \| null | settled cost; `null` for the brief window before settlement flushes (batched on a ~100 ms cadence) |
| `recent_attempts[].at` | string | dispatch timestamp |

Both audit endpoints also return the following per-attempt accounting fields:

| Field | Meaning |
|---|---|
| `billing_source` | `house` for marketplace-funded inference; `byok` for inference on your own provider connection |
| `inference_cost_usd` | Cost attributed from pinned rates and usage; `null` when either is unknown |
| `inference_price_source` | `platform_list`, `customer_rate`, `unknown`, or `legacy_platform_list` for historical records |
| `price_snapshot_id` | Identifier of the pinned rate snapshot; `null` when unknown or unavailable historically |
| `fee_policy_version` | Economic policy applied to this attempt: `house-pass-through-v1` or `byok-free-v1` |

`cost_usd` is the marketplace charge. `inference_cost_usd` is separate cost attribution:
it is not a provider invoice, and a list-price or estimated-usage calculation can differ
from the customer's actual provider bill. BYOK attempts have a zero marketplace
charge and an unknown (`null`) inference cost. Spend caps apply to marketplace charges.
See [Bring your own key](/docs/byok) and [routing policies](/docs/routing-policies).

Rows are **attempts**, not requests: a request that failed over appears once
per physical dispatch, sharing one `request_id`.

## GET /v1/generation?id=

The full per-attempt audit for one request. The `id` is the `x-request-id`
header every response carries (also `request_id` in `/v1/usage` rows).

| Parameter | In | Required | Meaning |
|---|---|---|---|
| `id` | query | yes | the request id to audit |

An `id` that is not a request id at all (not a UUID) is a `404 not_found`. An
unknown id — or one belonging to another org, workspace or principal — returns
`404` with `"attempts": []` and `"admission_events": []`. Here is a request
that failed over once (illustrative deployment ids; yours match your
`x-tm-provider` headers):

```bash
curl -s "https://api.routerplus.com/v1/generation?id=0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33" \
  -H "Authorization: Bearer $TM_API_KEY" | jq
```

```json
{
  "request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33",
  "admission_events": [],
  "attempts": [
    {
      "attempt": 1,
      "deployment": "provider-a",
      "model": "claude-sonnet-4-5",
      "outcome": "failed",
      "usage_provenance": "observed",
      "tokens": { "input": 1200, "cached": 0, "cache_write": 0, "output": 0, "reasoning": 0 },
      "reserved_max_usd": 0.009,
      "cost_usd": 0.0036,
      "billing_source": "house",
      "inference_cost_usd": 0.0036,
      "inference_price_source": "platform_list",
      "price_snapshot_id": "5d2c…",
      "fee_policy_version": "house-pass-through-v1",
      "error_code": "upstream_error",
      "error_origin": "upstream",
      "admission_context": { "…": "…" },
      "route_context": { "…": "…" },
      "dropped_params": [],
      "dispatched_at": "2026-09-04T09:12:31.512Z"
    },
    {
      "attempt": 2,
      "deployment": "provider-b",
      "model": "claude-sonnet-4-5",
      "outcome": "completed",
      "usage_provenance": "observed",
      "tokens": { "input": 1200, "cached": 800, "cache_write": 0, "output": 350, "reasoning": 0 },
      "reserved_max_usd": 0.009,
      "cost_usd": 0.00669,
      "billing_source": "house",
      "inference_cost_usd": 0.00669,
      "inference_price_source": "platform_list",
      "price_snapshot_id": "5d2c…",
      "fee_policy_version": "house-pass-through-v1",
      "error_code": "",
      "error_origin": null,
      "admission_context": { "…": "…" },
      "route_context": { "…": "…" },
      "dropped_params": [],
      "dispatched_at": "2026-09-04T09:12:33.104Z"
    }
  ]
}
```

| Field | Type | Meaning |
|---|---|---|
| `attempt` | number | physical dispatch ordinal, starting at 1 |
| `deployment` | string | deployment that served this attempt; `routerplus` on a closed dedicated endpoint |
| `outcome` | string | see below |
| `usage_provenance` | string | `observed` (provider-reported tokens, or a completed attempt), `unknown` (no usage ever seen), or `estimated` — the provider reported nothing and the gateway priced the attempt itself: on [`/v1/images/generations`](/docs/api-images) a 200 without usage bills `n` × the model's per-image ceiling, the amount the reservation held; on a chat attempt that streamed output, that output at ~4 characters per token |
| `tokens` | object | `input` (cache-inclusive), `cached` (reads), `cache_write`, `output`, `reasoning` (subset of output). On a model billed at the provider's reported cost, `output` is the charge in micro-dollars |
| `reserved_max_usd` | number | the worst-case hold taken **before** dispatch (ceiling math) |
| `cost_usd` | number \| null | what actually settled (floor math); `null` only pre-settlement |
| `error_code` | string | canonical error class on failure, empty on success |
| `error_origin` | string \| null | where a failure came from — `upstream`, `gateway_admission`, `upstream_quota`, `gateway_infrastructure` or `authorization`; `null` on success |
| `admission_context` | object \| null | the limits this attempt claimed, as recorded at dispatch — see [Limits and capacity](/docs/admission) |
| `route_context` | object \| null | the route plan behind this attempt: plan id, the deployments planned and excluded, the policy revisions in force; on a dedicated endpoint, `dedicated` (the endpoint, its revision and the role) |
| `dropped_params` | array | the parameter paths the gateway dropped for this deployment — the same names as the `x-tm-dropped-params` header |

On a closed [dedicated endpoint](/docs/dedicated-endpoints#closed-endpoints), `deployment`
and `provider` are `routerplus`, `model` is the endpoint id, `route_context` keeps only
`dedicated`, `admission_context` is null, `dropped_params` is empty, and `admission_events`
carry no pool `scope_id`. On any dedicated endpoint, its dedicated capacity reads as
`routerplus`.

`admission_events` lists the limit or balance refusals recorded for this request
before any attempt (`origin`, `scope`, `scope_id`, `limit_kind`, `reason`,
`status`, `occurred_at`); it is empty when nothing was refused.

`reserved_max_usd` versus `cost_usd` is the reserve-then-settle contract made
visible: the hold is deliberately pessimistic and is released in full; the
settle is computed from provider-reported usage and rounds down. Details in
[Pricing & billing](/docs/pricing).

### Attempt outcomes

| `outcome` | Meaning | Billing |
|---|---|---|
| `completed` | full response delivered | settled on reported usage |
| `failed` | attempt failed; a later attempt may have served you | settled on any usage the provider reported before dying (often 0) |
| `cancelled` | you disconnected mid-request | settled on usage streamed up to the disconnect; if you disconnected before the attempt left the gateway, it settles at zero with provenance `observed`. On `/v1/images/generations` the render is never aborted: it finishes upstream and the attempt settles the provider's reported usage |
| `incomplete` | provider died after your stream committed; you received a terminal error event. On `/v1/videos`, a job the gateway gave up following after two hours | settled on usage observed so far; a video job that timed out settles at its reservation |
| `unknown_after_crash` | the gateway could not observe the outcome (it stopped mid-attempt; the attempt is recovered from the gateway's durable ledger journal at restart, or by the operator reconciler) | settles **at the reservation** — the call may have been billed upstream, so the ledger keeps the conservative number |

> [!WARNING]
> You are debited for **every settled attempt** of a request, including ones
> discarded by pre-commit failover — an upstream that dies after reading your
> prompt has still billed prompt tokens. That is why the wire's `usage.cost`
> is the request total, and why this endpoint exists: the slices are never
> hidden. In the example above the wire reported `"cost": 0.01029` —
> `0.0036 + 0.00669`, both attempts, no averaging.

## Recomputing your bill from the wire

Sum of attempt costs equals the response's `usage.cost`:

```bash
curl -s "https://api.routerplus.com/v1/generation?id=$REQ" -H "Authorization: Bearer $TM_API_KEY" \
  | jq '[.attempts[].cost_usd] | add'
```

And each attempt's cost recomputes from its token counts and the published
prices — the exact settle math, floor division and all:

```python
import os
import requests
from decimal import Decimal

GATEWAY = "https://api.routerplus.com"
H = {"Authorization": f"Bearer {os.environ['TM_API_KEY']}"}

prices = {m["id"]: m["pricing_usd_per_million"]
          for m in requests.get("https://app.routerplus.com/api/models.json").json()["models"]}

def per_m(p, sku, fallback):
    # absent SKUs bill at the base rate — never $0
    return int(Decimal(p.get(sku, p[fallback])) * 1_000_000)  # micro-USD per M

def settle_micro(t, p):
    non_cached = max(0, t["input"] - t["cached"] - t["cache_write"])
    reasoning = min(max(0, t["reasoning"]), t["output"])
    plain_out = t["output"] - reasoning
    total = (non_cached        * per_m(p, "prompt", "prompt")
           + t["cached"]       * per_m(p, "cached_prompt", "prompt")
           + t["cache_write"]  * per_m(p, "cache_write", "prompt")
           + plain_out         * per_m(p, "completion", "completion")
           + reasoning         * per_m(p, "internal_reasoning", "completion"))
    return total // 1_000_000  # floor: settlement rounds down

gen = requests.get(f"{GATEWAY}/v1/generation",
                   params={"id": os.environ["REQ"]}, headers=H).json()
micro = sum(settle_micro(a["tokens"], prices[a["model"]])
            for a in gen["attempts"] if a["outcome"] != "unknown_after_crash")
print(micro / 1e6)  # equals the wire's usage.cost
```

(`unknown_after_crash` attempts are excluded because they settle at the
reservation, not from token counts — compare their `reserved_max_usd`. The same
math holds for a model billed at the provider's reported cost: its prices are
`"0"` in and `"1"` out per million, and `output` is the charge in
micro-dollars. An attempt with `inference_price_source` `customer_rate` was
billed at a [dedicated endpoint's](/docs/dedicated-endpoints#prices) contract
rate card: recompute it with those prices instead.)

One request re-priced at today's catalog can differ from what you were charged
only if the price changed since dispatch: attempts settle against **frozen
price snapshots** pinned when the request was reserved, which is the
recomputation guarantee's fine print — see
[Pricing & billing](/docs/pricing).

## Related

- [Pricing & billing](/docs/pricing) — the reserve/settle math these numbers come from
- [Errors & remediation](/docs/errors) — what each `error_code` means and what to do
- [Quickstart](/docs/quickstart) — signup to first billed call

## Who can read what

Usage and generation reads are restricted to the authenticated workspace and
principal, including that principal's other keys. A request made under another
workspace or principal is a 404, like one that never existed. Keys bound to an
explicit workspace or principal receive `null` for the org balance, credit and
spend totals; a key in the default workspace and principal gets them. Members
of the organization see the org-wide accounting in the console at
[https://app.routerplus.com/console/usage](https://app.routerplus.com/console/usage) and
[https://app.routerplus.com/console/logs](https://app.routerplus.com/console/logs).
