# Rate limits, capacity and outage behavior

Admission is shared across gateway instances: every limit on this page is counted in one store, so two instances never admit the same minute twice. This page is the full picture — scopes, provider pools, fairness, and what happens when a store is down. The everyday view (per-key RPM, the org defaults, spend caps) is on [Rate limits & spend caps](/docs/limits). BYOK connections and routing policies, which pools attach to, are on [BYOK setup](/docs/byok) and [routing policies](/docs/routing-policies).

## Manage limits

Sign in and open [`/console/limits`](https://app.routerplus.com/console/limits) to view or set gateway limits, create provider pools and
assign connections. Pool updates replace the supplied limits and parent; enter the complete
desired configuration. Gateway limits overlap: the strictest applicable platform, org,
workspace, principal, key, model and provider-account constraint applies.

Management APIs use the site origin (`https://app.routerplus.com`), session cookie, JSON content type and same-origin
mutation protection. Inference bearer keys cannot administer them. Owners and admins can
write; viewers can read.

| Endpoint | Purpose |
|---|---|
| `GET /api/admission-limits` | List your org's configured limits, every scope |
| `POST /api/admission-limits` | Replace one scope's limits |
| `GET /api/quota-pools` | List the newest 100 org-owned provider pools |
| `GET /api/quota-pools/{id}` | Read one pool |
| `POST /api/quota-pools` | Create a pool |
| `POST /api/quota-pools/{id}` | Replace its limits and parent |
| `POST /api/quota-pools/{id}/bind` | Assign an owned connection to that pool |

Example gateway limit body:

```json
{"scope":"key","subject":"VIRTUAL_KEY_UUID","rpm":100,"tpm":200000,"concurrency":4,"outage_mode":"bounded_open"}
```

`scope` is `org`, `workspace`, `principal`, `key` or `model`. For org scope, subject is `org`; for
workspace, principal and key scopes it is the UUID; for model scope it is
the exact model ID. Null/omitted numeric values inherit defaults or parent restrictions.
RPM accepts 1–1,000,000, TPM 1–1,000,000,000, concurrency 1–10,000. A key's original RPM
ceiling still applies, so configuring a larger key limit does not raise that ceiling.
Use `outage_mode:"closed"` for explicit refusal during counter-store failure; default is
`bounded_open`. A closed mode at any applicable scope wins.

Default limits are 300 RPM, 1M TPM and 8 concurrent requests for an org, a workspace
and a principal. Until an org has added credit, the org and each of its keys are held to
20 RPM (an operator override replaces this, as it replaces the defaults). Keys default to their RPM ceiling, 1M TPM and 8 concurrent requests.
A model scope has no default: set it to add a limit for one model. The operator
configures platform limits independently. Under a contract, the operator can also replace
the defaults and the keys' RPM ceiling for one org (an override, with a per-second limit
on the org if needed); limits you set still apply and can only lower them. A
[dedicated endpoint's](/docs/dedicated-endpoints#limits) traffic keeps the endpoint's own
limits.

## Declare provider account capacity

```json
{
  "provider":"openai",
  "rpm":500,
  "tpm":1000000,
  "concurrency":20,
  "headroom_percent":90,
  "parent_pool_id":null
}
```

Use `openai`, `anthropic`, `azure` or `bedrock` for customer pools. Supply finite positive values.
`headroom_percent` is the usable percentage, from 1–100; default 90. Fractional limits
round down, with a minimum of one unit. A pool parent must be owned by the same org and use
the same provider. Up to four acyclic levels are supported, including the leaf.

Bind with `{"connection_id":"CONNECTION_UUID"}` at `/api/quota-pools/{id}/bind`; the
connection's profile must match the pool's provider.
Keys for the same upstream account should share a pool, or child pools under a shared
account parent. Rotating credentials does not reset capacity. Parent and child claims are
checked together. The service cannot infer your real provider-account identity from a key;
accurate pool declarations remain your responsibility.

Marketplace house pools are operator-owned and mapped to catalog deployment IDs; customers
cannot change them. A route with no declared pool is unavailable (`503`, `x-tm-limit-kind:
capacity_unknown`), rather than implicitly
unlimited. Provider usage outside this gateway is not observed. Provider 429s cool the
affected leaf pool for the provider's `retry-after` (1 to 60 seconds), shared across
workers; an unknown error scope is not promoted to a global provider outage.

## Counting and token estimates

One admitted logical inference request consumes gateway RPM once. Every physical attempt,
including an eligible fallback, separately consumes upstream RPM and token reservation.
Requests already sent to a provider are not refunded as if they never happened. A request
proven not dispatched may release its provider claim. Concurrency uses renewable ten-second
leases so a crashed worker does not strand slots.

Input reservation uses UTF-8 byte size plus a conservative media allowance (65,536 per image
or document part). The base64 of an image sent inline is left out of the byte size: the
image counts its allowance only. It is an
estimate, not a vendor tokenizer guarantee. Requests enforce an output bound of 1–32,768
tokens, default 4,096, on the outgoing wire. Larger or conflicting output limits return 400.
Completed, observed input and output usage replaces the estimate; cache and reasoning
subsets are not counted twice. Partial prefill is not a final bill: cancelled, incomplete,
unknown or estimated usage keeps the reservation until its rate window expires, except
for proven never-dispatched work. The rolling-minute implementation includes a conservative partial
second, so a charge can remain for up to 61 seconds.

Anthropic token counting uses the same authorization and shared limits. It consumes gateway
and provider RPM, reserves estimated input capacity and records a durable request audit
before calling upstream. It does not create an inference debit or inference-attempt row.
Its answer is `{"input_tokens": N}` and nothing else.

## Inspect limits and errors

`GET /v1/limits?model=MODEL_ID` on the gateway returns effective gateway limits for the key,
including inherited restrictions, the output bound and outage mode. Model is optional.
Successful admission exposes `x-tm-remaining-rpm` and `x-tm-remaining-tpm`, the minimum
remaining amount across claimed scopes at that moment. These are snapshots, not promises
that capacity will still be available for a later request. On a dedicated endpoint with a
[paced limit](/docs/dedicated-endpoints#the-paced-limit), `x-tm-queue-ms` says how long the
request waited for its turn before admission.

Failures retain SDK-compatible statuses and include metadata and CORS-readable headers:

| Origin | Typical status | Meaning |
|---|---:|---|
| `gateway_admission` | 429 or 503 | Customer quota/spend limit, fairness headroom or worker capacity |
| `upstream_quota` | 429; 503 if unconfigured | Declared account pool exhaustion/cooldown or unknown capacity |
| `upstream` | Provider status | An actual provider response, including its 429 |
| `gateway_infrastructure` | 503 | Counter/journal/recovery failure |
| `authorization` | 401, 404 or 503 | Invalid/unavailable authority or changed/expired snapshot |

Headers include `x-tm-error-origin`, `x-tm-limit-scope`, `x-tm-limit-kind` (`rpm`, `rps`,
`rps_queue`, `tpm`, `concurrency`, `cooldown`, `spend`, `health`, `capacity_unknown`, …), optional
`x-tm-limit-id`, and `retry-after`. Metadata includes retryability and, where known, retry
or cap-reset time. A rolling-window retry time is advisory; it is not a capacity guarantee.

Quota/admission rejections after authentication are recorded even without a physical
attempt. The record is best effort: when it cannot be written, you still get the 429, not a
503. Query `GET /v1/generation?id=REQUEST_ID`: `admission_events` accompany
attempts, and attempts include receipt context/error origin. Unauthenticated requests and
HTTP-parser protections cannot be attributed to an authenticated org in this audit.

## Availability and fairness

- **Healthy shared store:** limits coordinate across workers. Pilot fairness protects idle
  tenant headroom and caps each org at a share of a shared house pool's usable capacity —
  half by default, 75% on the decision-model pools (`typesafe`, `openrouter-decisions`,
  `openrouter-decisions-free`, `bespokelabs`, `workers-ai`, `levanto`, `fastino`, `routerplus`, `perplexity-decisions`) — with a one-unit minimum. It rejects excess work immediately; there is no waiting queue, except on a dedicated endpoint with a
  [paced limit](/docs/dedicated-endpoints#the-paced-limit), and no general no-starvation
  guarantee for arbitrary tenant counts.
- **Redis unavailable:** the default bounded local allowance is at most five logical
  requests/minute, two concurrent local requests and 25,000 estimated TPM per scope per
  worker, for at most 30 seconds. A dedicated endpoint's traffic instead gets the
  endpoint's own limits divided among the gateways that serve it, for at most 10 minutes
  by default; a paced endpoint paces each gateway at its share of the rate.
  `x-tm-admission-mode: bounded_local` marks this path.
  Limits and balances can be approximate across workers during this interval. Strict
  mode refuses new work. Every provider call still requires a durable journal intent.
- **Postgres unavailable:** previously verified snapshots may serve for a bounded time
  (five minutes by default; the operator can set up to one hour for an outage) while the
  live Redis authority fence still matches. A single failed read falls back to a snapshot at
  most five minutes old. Once reads have failed for five seconds with none succeeding, the
  gateway answers from the snapshot at once, up to 15 seconds after the last failure,
  instead of waiting on Postgres again. Cold, evicted,
  expired or changed authority fails closed. `x-tm-authorization-mode: bounded_snapshot`
  identifies requests that used this fallback. If both authorities are unavailable, cached
  authority is not used.
- **Management during Redis failure:** mutations refuse rather than losing revocation
  coordination. A failed publication can leave a deliberate fence that needs operator
  recovery. Existing provider calls may finish on their pinned credential version.
- **Redis process replacement/lost state:** shared counters require explicit recovery;
  an empty store is not treated as a full unused allowance.

These are pilot controls. They do not establish provider capacity commitments, exact vendor
TPM compliance, private connectivity, regional processing or
ZDR. Under a routing policy, automatic fallback stays within one provider label (D10);
see [routing policies](/docs/routing-policies).

## Workspace and principal scopes

`workspace` and `principal` limits aggregate across the keys bound to that scope,
inside the same atomic claim as org/key/model/platform/pool constraints. Bindings come from
server-managed keys, never request headers. See [identity management](identity.md).
