Console
API reference/GET /v1/usage & /v1/generation

GET /v1/usage & /v1/generation

Spend, balances, and the per-attempt audit trail.

/llms.txt

Two read-only endpoints on the gateway answer "what have I spent?" and "what exactly did this request cost, attempt by attempt?". They read the same ledger your balance is derived from — there is no separate reporting pipeline that could disagree with your bill.

Both are metadata only: token counts, outcomes, and costs. Prompts and completions are never stored, so they cannot be returned.

Authentication

Both endpoints require a billed key (the first key shown after verified sign-in, or one created in the console) and are scoped to that key's org, workspace and principal. Send it either way:

bash
curl -s https://api.routerplus.com/v1/usage -H "Authorization: Bearer $TM_API_KEY"
curl -s https://api.routerplus.com/v1/usage -H "x-api-key: $TM_API_KEY"
Note

A static development key (a gateway running without the money plane) has no ledger to read, so on these routes it gets the same undifferentiated 404 as any unknown path.

GET /v1/usage

The 20 most recent attempts made with the keys of your workspace and principal, newest first. The org balance fields are filled for a key in the organization's default workspace and default principal — every first key shown after verified sign-in or made in the console — and null for a key bound to an explicit workspace or principal (see Workspaces and principals).

bash
curl -s https://api.routerplus.com/v1/usage -H "Authorization: Bearer $TM_API_KEY" | jq
json
{
  "balance_usd": 0.98971,
  "credited_usd": 1.0,
  "spent_usd": 0.01029,
  "recent_attempts": [
    {
      "request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33",
      "deployment": "provider-b",
      "model": "claude-sonnet-4-5",
      "outcome": "completed",
      "usage_provenance": "observed",
      "input_tokens": 1200,
      "output_tokens": 350,
      "cost_usd": 0.00669,
      "billing_source": "house",
      "inference_cost_usd": 0.00669,
      "inference_price_source": "platform_list",
      "price_snapshot_id": "5d2c…",
      "fee_policy_version": "house-pass-through-v1",
      "at": "2026-09-04T09:12:33.104Z"
    },
    {
      "request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33",
      "deployment": "provider-a",
      "model": "claude-sonnet-4-5",
      "outcome": "failed",
      "usage_provenance": "observed",
      "input_tokens": 1200,
      "output_tokens": 0,
      "cost_usd": 0.0036,
      "billing_source": "house",
      "inference_cost_usd": 0.0036,
      "inference_price_source": "platform_list",
      "price_snapshot_id": "5d2c…",
      "fee_policy_version": "house-pass-through-v1",
      "at": "2026-09-04T09:12:31.512Z"
    }
  ]
}
FieldTypeMeaning
balance_usdnumber | nullcredited_usd − spent_usd, derived from ledger rows on read — never a cached counter
credited_usdnumber | nullsum of all credit grants to the org
spent_usdnumber | nullsum of every settled attempt
recent_attempts[].request_idstringthe request's x-request-id; feed it to /v1/generation
recent_attempts[].deploymentstringwhich deployment served the attempt — matches the x-tm-provider response header; routerplus on a closed dedicated endpoint and on any endpoint's dedicated capacity
recent_attempts[].modelstringthe model id you requested; on a closed dedicated endpoint, the endpoint id
recent_attempts[].outcomestringsee the outcome table
recent_attempts[].usage_provenancestring | nullobserved, estimated or unknown; null before settlement
recent_attempts[].input_tokens / output_tokensnumberprovider-reported counts (input is cache-inclusive). On a model billed at the provider's reported cost, output_tokens is the charge in micro-dollars
recent_attempts[].cost_usdnumber | nullsettled cost; null for the brief window before settlement flushes (batched on a ~100 ms cadence)
recent_attempts[].atstringdispatch timestamp

Both audit endpoints also return the following per-attempt accounting fields:

FieldMeaning
billing_sourcehouse for marketplace-funded inference; byok for inference on your own provider connection
inference_cost_usdCost attributed from pinned rates and usage; null when either is unknown
inference_price_sourceplatform_list, customer_rate, unknown, or legacy_platform_list for historical records
price_snapshot_idIdentifier of the pinned rate snapshot; null when unknown or unavailable historically
fee_policy_versionEconomic policy applied to this attempt: house-pass-through-v1 or byok-free-v1

cost_usd is the marketplace charge. inference_cost_usd is separate cost attribution: it is not a provider invoice, and a list-price or estimated-usage calculation can differ from the customer's actual provider bill. BYOK attempts have a zero marketplace charge and an unknown (null) inference cost. Spend caps apply to marketplace charges. See Bring your own key and routing policies.

Rows are attempts, not requests: a request that failed over appears once per physical dispatch, sharing one request_id.

GET /v1/generation?id=

The full per-attempt audit for one request. The id is the x-request-id header every response carries (also request_id in /v1/usage rows).

ParameterInRequiredMeaning
idqueryyesthe request id to audit

An id that is not a request id at all (not a UUID) is a 404 not_found. An unknown id — or one belonging to another org, workspace or principal — returns 404 with "attempts": [] and "admission_events": []. Here is a request that failed over once (illustrative deployment ids; yours match your x-tm-provider headers):

bash
curl -s "https://api.routerplus.com/v1/generation?id=0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33" \
  -H "Authorization: Bearer $TM_API_KEY" | jq
json
{
  "request_id": "0b6f6a4e-6a01-4c5e-9d5a-1c9f2b7e8a33",
  "admission_events": [],
  "attempts": [
    {
      "attempt": 1,
      "deployment": "provider-a",
      "model": "claude-sonnet-4-5",
      "outcome": "failed",
      "usage_provenance": "observed",
      "tokens": { "input": 1200, "cached": 0, "cache_write": 0, "output": 0, "reasoning": 0 },
      "reserved_max_usd": 0.009,
      "cost_usd": 0.0036,
      "billing_source": "house",
      "inference_cost_usd": 0.0036,
      "inference_price_source": "platform_list",
      "price_snapshot_id": "5d2c…",
      "fee_policy_version": "house-pass-through-v1",
      "error_code": "upstream_error",
      "error_origin": "upstream",
      "admission_context": { "…": "…" },
      "route_context": { "…": "…" },
      "dropped_params": [],
      "dispatched_at": "2026-09-04T09:12:31.512Z"
    },
    {
      "attempt": 2,
      "deployment": "provider-b",
      "model": "claude-sonnet-4-5",
      "outcome": "completed",
      "usage_provenance": "observed",
      "tokens": { "input": 1200, "cached": 800, "cache_write": 0, "output": 350, "reasoning": 0 },
      "reserved_max_usd": 0.009,
      "cost_usd": 0.00669,
      "billing_source": "house",
      "inference_cost_usd": 0.00669,
      "inference_price_source": "platform_list",
      "price_snapshot_id": "5d2c…",
      "fee_policy_version": "house-pass-through-v1",
      "error_code": "",
      "error_origin": null,
      "admission_context": { "…": "…" },
      "route_context": { "…": "…" },
      "dropped_params": [],
      "dispatched_at": "2026-09-04T09:12:33.104Z"
    }
  ]
}
FieldTypeMeaning
attemptnumberphysical dispatch ordinal, starting at 1
deploymentstringdeployment that served this attempt; routerplus on a closed dedicated endpoint
outcomestringsee below
usage_provenancestringobserved (provider-reported tokens, or a completed attempt), unknown (no usage ever seen), or estimated — the provider reported nothing and the gateway priced the attempt itself: on /v1/images/generations a 200 without usage bills n × the model's per-image ceiling, the amount the reservation held; on a chat attempt that streamed output, that output at ~4 characters per token
tokensobjectinput (cache-inclusive), cached (reads), cache_write, output, reasoning (subset of output). On a model billed at the provider's reported cost, output is the charge in micro-dollars
reserved_max_usdnumberthe worst-case hold taken before dispatch (ceiling math)
cost_usdnumber | nullwhat actually settled (floor math); null only pre-settlement
error_codestringcanonical error class on failure, empty on success
error_originstring | nullwhere a failure came from — upstream, gateway_admission, upstream_quota, gateway_infrastructure or authorization; null on success
admission_contextobject | nullthe limits this attempt claimed, as recorded at dispatch — see Limits and capacity
route_contextobject | nullthe route plan behind this attempt: plan id, the deployments planned and excluded, the policy revisions in force; on a dedicated endpoint, dedicated (the endpoint, its revision and the role)
dropped_paramsarraythe parameter paths the gateway dropped for this deployment — the same names as the x-tm-dropped-params header

On a closed dedicated endpoint, deployment and provider are routerplus, model is the endpoint id, route_context keeps only dedicated, admission_context is null, dropped_params is empty, and admission_events carry no pool scope_id. On any dedicated endpoint, its dedicated capacity reads as routerplus.

admission_events lists the limit or balance refusals recorded for this request before any attempt (origin, scope, scope_id, limit_kind, reason, status, occurred_at); it is empty when nothing was refused.

reserved_max_usd versus cost_usd is the reserve-then-settle contract made visible: the hold is deliberately pessimistic and is released in full; the settle is computed from provider-reported usage and rounds down. Details in Pricing & billing.

Attempt outcomes

outcomeMeaningBilling
completedfull response deliveredsettled on reported usage
failedattempt failed; a later attempt may have served yousettled on any usage the provider reported before dying (often 0)
cancelledyou disconnected mid-requestsettled on usage streamed up to the disconnect; if you disconnected before the attempt left the gateway, it settles at zero with provenance observed. On /v1/images/generations the render is never aborted: it finishes upstream and the attempt settles the provider's reported usage
incompleteprovider died after your stream committed; you received a terminal error event. On /v1/videos, a job the gateway gave up following after two hourssettled on usage observed so far; a video job that timed out settles at its reservation
unknown_after_crashthe gateway could not observe the outcome (it stopped mid-attempt; the attempt is recovered from the gateway's durable ledger journal at restart, or by the operator reconciler)settles at the reservation — the call may have been billed upstream, so the ledger keeps the conservative number
Warning

You are debited for every settled attempt of a request, including ones discarded by pre-commit failover — an upstream that dies after reading your prompt has still billed prompt tokens. That is why the wire's usage.cost is the request total, and why this endpoint exists: the slices are never hidden. In the example above the wire reported "cost": 0.01029 — 0.0036 + 0.00669, both attempts, no averaging.

Recomputing your bill from the wire

Sum of attempt costs equals the response's usage.cost:

bash
curl -s "https://api.routerplus.com/v1/generation?id=$REQ" -H "Authorization: Bearer $TM_API_KEY" \
  | jq '[.attempts[].cost_usd] | add'

And each attempt's cost recomputes from its token counts and the published prices — the exact settle math, floor division and all:

python
import os
import requests
from decimal import Decimal

GATEWAY = "https://api.routerplus.com"
H = {"Authorization": f"Bearer {os.environ['TM_API_KEY']}"}

prices = {m["id"]: m["pricing_usd_per_million"]
          for m in requests.get("https://app.routerplus.com/api/models.json").json()["models"]}

def per_m(p, sku, fallback):
    # absent SKUs bill at the base rate — never $0
    return int(Decimal(p.get(sku, p[fallback])) * 1_000_000)  # micro-USD per M

def settle_micro(t, p):
    non_cached = max(0, t["input"] - t["cached"] - t["cache_write"])
    reasoning = min(max(0, t["reasoning"]), t["output"])
    plain_out = t["output"] - reasoning
    total = (non_cached        * per_m(p, "prompt", "prompt")
           + t["cached"]       * per_m(p, "cached_prompt", "prompt")
           + t["cache_write"]  * per_m(p, "cache_write", "prompt")
           + plain_out         * per_m(p, "completion", "completion")
           + reasoning         * per_m(p, "internal_reasoning", "completion"))
    return total // 1_000_000  # floor: settlement rounds down

gen = requests.get(f"{GATEWAY}/v1/generation",
                   params={"id": os.environ["REQ"]}, headers=H).json()
micro = sum(settle_micro(a["tokens"], prices[a["model"]])
            for a in gen["attempts"] if a["outcome"] != "unknown_after_crash")
print(micro / 1e6)  # equals the wire's usage.cost

(unknown_after_crash attempts are excluded because they settle at the reservation, not from token counts — compare their reserved_max_usd. The same math holds for a model billed at the provider's reported cost: its prices are "0" in and "1" out per million, and output is the charge in micro-dollars. An attempt with inference_price_source customer_rate was billed at a dedicated endpoint's contract rate card: recompute it with those prices instead.)

One request re-priced at today's catalog can differ from what you were charged only if the price changed since dispatch: attempts settle against frozen price snapshots pinned when the request was reserved, which is the recomputation guarantee's fine print — see Pricing & billing.

Who can read what

Usage and generation reads are restricted to the authenticated workspace and principal, including that principal's other keys. A request made under another workspace or principal is a 404, like one that never existed. Keys bound to an explicit workspace or principal receive null for the org balance, credit and spend totals; a key in the default workspace and principal gets them. Members of the organization see the org-wide accounting in the console at https://app.routerplus.com/console/usage and https://app.routerplus.com/console/logs.

Markdown source for agents: /docs/api-usage.md · index at /llms.txt