GET /v1/usage & /v1/generation
Spend, balances, and the per-attempt audit trail.
Two read-only endpoints on the gateway answer "what have I spent?" and "what exactly did this request cost, attempt by attempt?". They read the same ledger your balance is derived from — there is no separate reporting pipeline that could disagree with your bill.
Both are metadata only: token counts, outcomes, and costs. Prompts and completions are never stored, so they cannot be returned.
Authentication
Both endpoints require a billed key (the first key shown after verified sign-in, or one created in the console) and are scoped to that key's org, workspace and principal. Send it either way:
curl -s https://api.routerplus.com/v1/usage -H "Authorization: Bearer $TM_API_KEY"
curl -s https://api.routerplus.com/v1/usage -H "x-api-key: $TM_API_KEY"A static development key (a gateway running without the money plane) has no ledger to read, so on these routes it gets the same undifferentiated 404 as any unknown path.
GET /v1/usage
The 20 most recent attempts made with the keys of your workspace and principal, newest first. The org balance fields are filled for a key in the organization's default workspace and default principal — every first key shown after verified sign-in or made in the console — and null for a key bound to an explicit workspace or principal (see Workspaces and principals).
curl -s https://api.routerplus.com/v1/usage -H "Authorization: Bearer $TM_API_KEY" | jq{
"balance_usd": 0.98971,
"credited_usd": 1.0,
"spent_usd": 0.01029,
"recent_attempts": [
{
"request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33",
"deployment": "provider-b",
"model": "claude-sonnet-4-5",
"outcome": "completed",
"usage_provenance": "observed",
"input_tokens": 1200,
"output_tokens": 350,
"cost_usd": 0.00669,
"billing_source": "house",
"inference_cost_usd": 0.00669,
"inference_price_source": "platform_list",
"price_snapshot_id": "5d2c…",
"fee_policy_version": "house-pass-through-v1",
"at": "2026-09-04T09:12:33.104Z"
},
{
"request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33",
"deployment": "provider-a",
"model": "claude-sonnet-4-5",
"outcome": "failed",
"usage_provenance": "observed",
"input_tokens": 1200,
"output_tokens": 0,
"cost_usd": 0.0036,
"billing_source": "house",
"inference_cost_usd": 0.0036,
"inference_price_source": "platform_list",
"price_snapshot_id": "5d2c…",
"fee_policy_version": "house-pass-through-v1",
"at": "2026-09-04T09:12:31.512Z"
}
]
}| Field | Type | Meaning |
|---|---|---|
balance_usd | number | null | credited_usd − spent_usd, derived from ledger rows on read — never a cached counter |
credited_usd | number | null | sum of all credit grants to the org |
spent_usd | number | null | sum of every settled attempt |
recent_attempts[].request_id | string | the request's x-request-id; feed it to /v1/generation |
recent_attempts[].deployment | string | which deployment served the attempt — matches the x-tm-provider response header; routerplus on a closed dedicated endpoint and on any endpoint's dedicated capacity |
recent_attempts[].model | string | the model id you requested; on a closed dedicated endpoint, the endpoint id |
recent_attempts[].outcome | string | see the outcome table |
recent_attempts[].usage_provenance | string | null | observed, estimated or unknown; null before settlement |
recent_attempts[].input_tokens / output_tokens | number | provider-reported counts (input is cache-inclusive). On a model billed at the provider's reported cost, output_tokens is the charge in micro-dollars |
recent_attempts[].cost_usd | number | null | settled cost; null for the brief window before settlement flushes (batched on a ~100 ms cadence) |
recent_attempts[].at | string | dispatch timestamp |
Both audit endpoints also return the following per-attempt accounting fields:
| Field | Meaning |
|---|---|
billing_source | house for marketplace-funded inference; byok for inference on your own provider connection |
inference_cost_usd | Cost attributed from pinned rates and usage; null when either is unknown |
inference_price_source | platform_list, customer_rate, unknown, or legacy_platform_list for historical records |
price_snapshot_id | Identifier of the pinned rate snapshot; null when unknown or unavailable historically |
fee_policy_version | Economic policy applied to this attempt: house-pass-through-v1 or byok-free-v1 |
cost_usd is the marketplace charge. inference_cost_usd is separate cost attribution: it is not a provider invoice, and a list-price or estimated-usage calculation can differ from the customer's actual provider bill. BYOK attempts have a zero marketplace charge and an unknown (null) inference cost. Spend caps apply to marketplace charges. See Bring your own key and routing policies.
Rows are attempts, not requests: a request that failed over appears once per physical dispatch, sharing one request_id.
GET /v1/generation?id=
The full per-attempt audit for one request. The id is the x-request-id header every response carries (also request_id in /v1/usage rows).
| Parameter | In | Required | Meaning |
|---|---|---|---|
id | query | yes | the request id to audit |
An id that is not a request id at all (not a UUID) is a 404 not_found. An unknown id — or one belonging to another org, workspace or principal — returns 404 with "attempts": [] and "admission_events": []. Here is a request that failed over once (illustrative deployment ids; yours match your x-tm-provider headers):
curl -s "https://api.routerplus.com/v1/generation?id=0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33" \
-H "Authorization: Bearer $TM_API_KEY" | jq{
"request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33",
"admission_events": [],
"attempts": [
{
"attempt": 1,
"deployment": "provider-a",
"model": "claude-sonnet-4-5",
"outcome": "failed",
"usage_provenance": "observed",
"tokens": { "input": 1200, "cached": 0, "cache_write": 0, "output": 0, "reasoning": 0 },
"reserved_max_usd": 0.009,
"cost_usd": 0.0036,
"billing_source": "house",
"inference_cost_usd": 0.0036,
"inference_price_source": "platform_list",
"price_snapshot_id": "5d2c…",
"fee_policy_version": "house-pass-through-v1",
"error_code": "upstream_error",
"error_origin": "upstream",
"admission_context": { "…": "…" },
"route_context": { "…": "…" },
"dropped_params": [],
"dispatched_at": "2026-09-04T09:12:31.512Z"
},
{
"attempt": 2,
"deployment": "provider-b",
"model": "claude-sonnet-4-5",
"outcome": "completed",
"usage_provenance": "observed",
"tokens": { "input": 1200, "cached": 800, "cache_write": 0, "output": 350, "reasoning": 0 },
"reserved_max_usd": 0.009,
"cost_usd": 0.00669,
"billing_source": "house",
"inference_cost_usd": 0.00669,
"inference_price_source": "platform_list",
"price_snapshot_id": "5d2c…",
"fee_policy_version": "house-pass-through-v1",
"error_code": "",
"error_origin": null,
"admission_context": { "…": "…" },
"route_context": { "…": "…" },
"dropped_params": [],
"dispatched_at": "2026-09-04T09:12:33.104Z"
}
]
}| Field | Type | Meaning |
|---|---|---|
attempt | number | physical dispatch ordinal, starting at 1 |
deployment | string | deployment that served this attempt; routerplus on a closed dedicated endpoint |
outcome | string | see below |
usage_provenance | string | observed (provider-reported tokens, or a completed attempt), unknown (no usage ever seen), or estimated — the provider reported nothing and the gateway priced the attempt itself: on /v1/images/generations a 200 without usage bills n × the model's per-image ceiling, the amount the reservation held; on a chat attempt that streamed output, that output at ~4 characters per token |
tokens | object | input (cache-inclusive), cached (reads), cache_write, output, reasoning (subset of output). On a model billed at the provider's reported cost, output is the charge in micro-dollars |
reserved_max_usd | number | the worst-case hold taken before dispatch (ceiling math) |
cost_usd | number | null | what actually settled (floor math); null only pre-settlement |
error_code | string | canonical error class on failure, empty on success |
error_origin | string | null | where a failure came from — upstream, gateway_admission, upstream_quota, gateway_infrastructure or authorization; null on success |
admission_context | object | null | the limits this attempt claimed, as recorded at dispatch — see Limits and capacity |
route_context | object | null | the route plan behind this attempt: plan id, the deployments planned and excluded, the policy revisions in force; on a dedicated endpoint, dedicated (the endpoint, its revision and the role) |
dropped_params | array | the parameter paths the gateway dropped for this deployment — the same names as the x-tm-dropped-params header |
On a closed dedicated endpoint, deployment and provider are routerplus, model is the endpoint id, route_context keeps only dedicated, admission_context is null, dropped_params is empty, and admission_events carry no pool scope_id. On any dedicated endpoint, its dedicated capacity reads as routerplus.
admission_events lists the limit or balance refusals recorded for this request before any attempt (origin, scope, scope_id, limit_kind, reason, status, occurred_at); it is empty when nothing was refused.
reserved_max_usd versus cost_usd is the reserve-then-settle contract made visible: the hold is deliberately pessimistic and is released in full; the settle is computed from provider-reported usage and rounds down. Details in Pricing & billing.
Attempt outcomes
outcome | Meaning | Billing |
|---|---|---|
completed | full response delivered | settled on reported usage |
failed | attempt failed; a later attempt may have served you | settled on any usage the provider reported before dying (often 0) |
cancelled | you disconnected mid-request | settled on usage streamed up to the disconnect; if you disconnected before the attempt left the gateway, it settles at zero with provenance observed. On /v1/images/generations the render is never aborted: it finishes upstream and the attempt settles the provider's reported usage |
incomplete | provider died after your stream committed; you received a terminal error event. On /v1/videos, a job the gateway gave up following after two hours | settled on usage observed so far; a video job that timed out settles at its reservation |
unknown_after_crash | the gateway could not observe the outcome (it stopped mid-attempt; the attempt is recovered from the gateway's durable ledger journal at restart, or by the operator reconciler) | settles at the reservation — the call may have been billed upstream, so the ledger keeps the conservative number |
You are debited for every settled attempt of a request, including ones discarded by pre-commit failover — an upstream that dies after reading your prompt has still billed prompt tokens. That is why the wire's usage.cost is the request total, and why this endpoint exists: the slices are never hidden. In the example above the wire reported "cost": 0.01029 — 0.0036 + 0.00669, both attempts, no averaging.
Recomputing your bill from the wire
Sum of attempt costs equals the response's usage.cost:
curl -s "https://api.routerplus.com/v1/generation?id=$REQ" -H "Authorization: Bearer $TM_API_KEY" \
| jq '[.attempts[].cost_usd] | add'And each attempt's cost recomputes from its token counts and the published prices — the exact settle math, floor division and all:
import os
import requests
from decimal import Decimal
GATEWAY = "https://api.routerplus.com"
H = {"Authorization": f"Bearer {os.environ['TM_API_KEY']}"}
prices = {m["id"]: m["pricing_usd_per_million"]
for m in requests.get("https://app.routerplus.com/api/models.json").json()["models"]}
def per_m(p, sku, fallback):
# absent SKUs bill at the base rate — never $0
return int(Decimal(p.get(sku, p[fallback])) * 1_000_000) # micro-USD per M
def settle_micro(t, p):
non_cached = max(0, t["input"] - t["cached"] - t["cache_write"])
reasoning = min(max(0, t["reasoning"]), t["output"])
plain_out = t["output"] - reasoning
total = (non_cached * per_m(p, "prompt", "prompt")
+ t["cached"] * per_m(p, "cached_prompt", "prompt")
+ t["cache_write"] * per_m(p, "cache_write", "prompt")
+ plain_out * per_m(p, "completion", "completion")
+ reasoning * per_m(p, "internal_reasoning", "completion"))
return total // 1_000_000 # floor: settlement rounds down
gen = requests.get(f"{GATEWAY}/v1/generation",
params={"id": os.environ["REQ"]}, headers=H).json()
micro = sum(settle_micro(a["tokens"], prices[a["model"]])
for a in gen["attempts"] if a["outcome"] != "unknown_after_crash")
print(micro / 1e6) # equals the wire's usage.cost(unknown_after_crash attempts are excluded because they settle at the reservation, not from token counts — compare their reserved_max_usd. The same math holds for a model billed at the provider's reported cost: its prices are "0" in and "1" out per million, and output is the charge in micro-dollars. An attempt with inference_price_source customer_rate was billed at a dedicated endpoint's contract rate card: recompute it with those prices instead.)
One request re-priced at today's catalog can differ from what you were charged only if the price changed since dispatch: attempts settle against frozen price snapshots pinned when the request was reserved, which is the recomputation guarantee's fine print — see Pricing & billing.
Related
error_code means and what to do →QuickstartSignup to first billed call →Who can read what
Usage and generation reads are restricted to the authenticated workspace and principal, including that principal's other keys. A request made under another workspace or principal is a 404, like one that never existed. Keys bound to an explicit workspace or principal receive null for the org balance, credit and spend totals; a key in the default workspace and principal gets them. Members of the organization see the org-wide accounting in the console at https://app.routerplus.com/console/usage and https://app.routerplus.com/console/logs.
Markdown source for agents: /docs/api-usage.md · index at /llms.txt