Console
Get started/Limits and capacity

Rate limits, capacity and outage behavior

Shared quotas, provider pools, fairness, and outage behavior.

/llms.txt

Admission is shared across gateway instances: every limit on this page is counted in one store, so two instances never admit the same minute twice. This page is the full picture — scopes, provider pools, fairness, and what happens when a store is down. The everyday view (per-key RPM, the org defaults, spend caps) is on Rate limits & spend caps. BYOK connections and routing policies, which pools attach to, are on BYOK setup and routing policies.

Manage limits

Sign in and open /console/limits to view or set gateway limits, create provider pools and assign connections. Pool updates replace the supplied limits and parent; enter the complete desired configuration. Gateway limits overlap: the strictest applicable platform, org, workspace, principal, key, model and provider-account constraint applies.

Management APIs use the site origin (https://app.routerplus.com), session cookie, JSON content type and same-origin mutation protection. Inference bearer keys cannot administer them. Owners and admins can write; viewers can read.

EndpointPurpose
GET /api/admission-limitsList your org's configured limits, every scope
POST /api/admission-limitsReplace one scope's limits
GET /api/quota-poolsList the newest 100 org-owned provider pools
GET /api/quota-pools/{id}Read one pool
POST /api/quota-poolsCreate a pool
POST /api/quota-pools/{id}Replace its limits and parent
POST /api/quota-pools/{id}/bindAssign an owned connection to that pool

Example gateway limit body:

json
{"scope":"key","subject":"VIRTUAL_KEY_UUID","rpm":100,"tpm":200000,"concurrency":4,"outage_mode":"bounded_open"}

scope is org, workspace, principal, key or model. For org scope, subject is org; for workspace, principal and key scopes it is the UUID; for model scope it is the exact model ID. Null/omitted numeric values inherit defaults or parent restrictions. RPM accepts 1–1,000,000, TPM 1–1,000,000,000, concurrency 1–10,000. A key's original RPM ceiling still applies, so configuring a larger key limit does not raise that ceiling. Use outage_mode:"closed" for explicit refusal during counter-store failure; default is bounded_open. A closed mode at any applicable scope wins.

Default limits are 300 RPM, 1M TPM and 8 concurrent requests for an org, a workspace and a principal. Until an org has added credit, the org and each of its keys are held to 20 RPM (an operator override replaces this, as it replaces the defaults). Keys default to their RPM ceiling, 1M TPM and 8 concurrent requests. A model scope has no default: set it to add a limit for one model. The operator configures platform limits independently. Under a contract, the operator can also replace the defaults and the keys' RPM ceiling for one org (an override, with a per-second limit on the org if needed); limits you set still apply and can only lower them. A dedicated endpoint's traffic keeps the endpoint's own limits.

Declare provider account capacity

json
{
  "provider":"openai",
  "rpm":500,
  "tpm":1000000,
  "concurrency":20,
  "headroom_percent":90,
  "parent_pool_id":null
}

Use openai, anthropic, azure or bedrock for customer pools. Supply finite positive values. headroom_percent is the usable percentage, from 1–100; default 90. Fractional limits round down, with a minimum of one unit. A pool parent must be owned by the same org and use the same provider. Up to four acyclic levels are supported, including the leaf.

Bind with {"connection_id":"CONNECTION_UUID"} at /api/quota-pools/{id}/bind; the connection's profile must match the pool's provider. Keys for the same upstream account should share a pool, or child pools under a shared account parent. Rotating credentials does not reset capacity. Parent and child claims are checked together. The service cannot infer your real provider-account identity from a key; accurate pool declarations remain your responsibility.

Marketplace house pools are operator-owned and mapped to catalog deployment IDs; customers cannot change them. A route with no declared pool is unavailable (503, x-tm-limit-kind: capacity_unknown), rather than implicitly unlimited. Provider usage outside this gateway is not observed. Provider 429s cool the affected leaf pool for the provider's retry-after (1 to 60 seconds), shared across workers; an unknown error scope is not promoted to a global provider outage.

Counting and token estimates

One admitted logical inference request consumes gateway RPM once. Every physical attempt, including an eligible fallback, separately consumes upstream RPM and token reservation. Requests already sent to a provider are not refunded as if they never happened. A request proven not dispatched may release its provider claim. Concurrency uses renewable ten-second leases so a crashed worker does not strand slots.

Input reservation uses UTF-8 byte size plus a conservative media allowance (65,536 per image or document part). The base64 of an image sent inline is left out of the byte size: the image counts its allowance only. It is an estimate, not a vendor tokenizer guarantee. Requests enforce an output bound of 1–32,768 tokens, default 4,096, on the outgoing wire. Larger or conflicting output limits return 400. Completed, observed input and output usage replaces the estimate; cache and reasoning subsets are not counted twice. Partial prefill is not a final bill: cancelled, incomplete, unknown or estimated usage keeps the reservation until its rate window expires, except for proven never-dispatched work. The rolling-minute implementation includes a conservative partial second, so a charge can remain for up to 61 seconds.

Anthropic token counting uses the same authorization and shared limits. It consumes gateway and provider RPM, reserves estimated input capacity and records a durable request audit before calling upstream. It does not create an inference debit or inference-attempt row. Its answer is {"input_tokens": N} and nothing else.

Inspect limits and errors

GET /v1/limits?model=MODEL_ID on the gateway returns effective gateway limits for the key, including inherited restrictions, the output bound and outage mode. Model is optional. Successful admission exposes x-tm-remaining-rpm and x-tm-remaining-tpm, the minimum remaining amount across claimed scopes at that moment. These are snapshots, not promises that capacity will still be available for a later request. On a dedicated endpoint with a paced limit, x-tm-queue-ms says how long the request waited for its turn before admission.

Failures retain SDK-compatible statuses and include metadata and CORS-readable headers:

OriginTypical statusMeaning
gateway_admission429 or 503Customer quota/spend limit, fairness headroom or worker capacity
upstream_quota429; 503 if unconfiguredDeclared account pool exhaustion/cooldown or unknown capacity
upstreamProvider statusAn actual provider response, including its 429
gateway_infrastructure503Counter/journal/recovery failure
authorization401, 404 or 503Invalid/unavailable authority or changed/expired snapshot

Headers include x-tm-error-origin, x-tm-limit-scope, x-tm-limit-kind (rpm, rps, rps_queue, tpm, concurrency, cooldown, spend, health, capacity_unknown, …), optional x-tm-limit-id, and retry-after. Metadata includes retryability and, where known, retry or cap-reset time. A rolling-window retry time is advisory; it is not a capacity guarantee.

Quota/admission rejections after authentication are recorded even without a physical attempt. The record is best effort: when it cannot be written, you still get the 429, not a

  1. 503
    Query GET /v1/generation?id=REQUEST_ID: admission_events accompany attempts, and attempts include receipt context/error origin. Unauthenticated requests and HTTP-parser protections cannot be attributed to an authenticated org in this audit.

Availability and fairness

  • Healthy shared store: limits coordinate across workers. Pilot fairness protects idle tenant headroom and caps each org at a share of a shared house pool's usable capacity — half by default, 75% on the decision-model pools (typesafe, openrouter-decisions, openrouter-decisions-free, bespokelabs, workers-ai, levanto, fastino, routerplus, perplexity-decisions) — with a one-unit minimum. It rejects excess work immediately; there is no waiting queue, except on a dedicated endpoint with a paced limit, and no general no-starvation guarantee for arbitrary tenant counts.
  • Redis unavailable: the default bounded local allowance is at most five logical requests/minute, two concurrent local requests and 25,000 estimated TPM per scope per worker, for at most 30 seconds. A dedicated endpoint's traffic instead gets the endpoint's own limits divided among the gateways that serve it, for at most 10 minutes by default; a paced endpoint paces each gateway at its share of the rate. x-tm-admission-mode: bounded_local marks this path. Limits and balances can be approximate across workers during this interval. Strict mode refuses new work. Every provider call still requires a durable journal intent.
  • Postgres unavailable: previously verified snapshots may serve for a bounded time (five minutes by default; the operator can set up to one hour for an outage) while the live Redis authority fence still matches. A single failed read falls back to a snapshot at most five minutes old. Once reads have failed for five seconds with none succeeding, the gateway answers from the snapshot at once, up to 15 seconds after the last failure, instead of waiting on Postgres again. Cold, evicted, expired or changed authority fails closed. x-tm-authorization-mode: bounded_snapshot identifies requests that used this fallback. If both authorities are unavailable, cached authority is not used.
  • Management during Redis failure: mutations refuse rather than losing revocation coordination. A failed publication can leave a deliberate fence that needs operator recovery. Existing provider calls may finish on their pinned credential version.
  • Redis process replacement/lost state: shared counters require explicit recovery; an empty store is not treated as a full unused allowance.

These are pilot controls. They do not establish provider capacity commitments, exact vendor TPM compliance, private connectivity, regional processing or ZDR. Under a routing policy, automatic fallback stays within one provider label (D10); see routing policies.

Workspace and principal scopes

workspace and principal limits aggregate across the keys bound to that scope, inside the same atomic claim as org/key/model/platform/pool constraints. Bindings come from server-managed keys, never request headers. See identity management.

Markdown source for agents: /docs/admission.md · index at /llms.txt