Console
Core concepts/Rate limits & spend caps

Rate limits & spend caps

Per-key RPM, organization token and concurrency limits, monthly caps, the 429 with a reset time.

/llms.txt

Several independent brakes can stop a request before it reaches a provider: the key's requests-per-minute limit, the organization's token and concurrency limits, the provider pool's capacity, the monthly spend caps, and the balance itself. Each refuses with a typed error that names its class in the body and in x-tm-error-code — nothing is throttled or queued silently, and a refused request costs nothing. The error shapes follow the same conventions as everything else — see Errors. Pools, fairness and what happens when the shared store is down are on Limits and capacity.

Per-key requests per minute

Every key has an RPM limit, enforced over a rolling 60-second window per key. The window lives in a shared store, so every gateway instance counts the same requests, and a request's charge can stay in the window for up to 61 seconds.

KeyRPM
First trial key shown after verified browser sign-in20
Key created in the console or with the identity API300
The console's managed Playground key60

Until the organization has bought credit (or holds credit we granted), every customer key and the organization as a whole run at the trial rate: 20 requests per minute, however many keys it creates. The first purchase lifts it within seconds; no key changes. Requests served only by your own provider keys (BYOK) and the managed Playground key are not held to it. Historical card-verification bonuses remain promotional credit, not purchases.

An organization on trial credit can also be limited when we detect unusual usage. Such a request gets 429 usage_limited (with retry-after) or 403 usage_restricted, and the message asks you to contact support, who can review the account. A purchase usually lifts it.

Over the limit, the request is refused before any reservation or provider work:

HTTP/1.1 429 Too Many Requests
retry-after: 1
x-tm-error-code: rate_limit
x-tm-error-origin: gateway_admission
x-tm-limit-scope: key
x-tm-limit-kind: rpm
x-tm-limit-id: key:<key id>
json
{
  "error": {
    "code": "rate_limit",
    "message": "configured quota exhausted",
    "type": "invalid_request_error",
    "metadata": {
      "error_type": "rate_limit",
      "origin": "gateway_admission",
      "limit_scope": "key",
      "limit_kind": "rpm",
      "limit_id": "key:<key id>",
      "retryable": true,
      "retry_at": "2026-09-24T10:14:04.000Z"
    }
  },
  "request_id": "..."
}

On the Anthropic surface the same refusal is a native rate_limit_error carrying the same error_type and metadata, so the Anthropic SDK's own backoff logic engages.

retry-after on a rolling window is advisory: the window frees one second at a time, and other requests may take the room first. Honor it and retry; don't hammer.

Note

The RPM values are fixed today: there is no self-serve way to raise a key's ceiling. On /console/limits you can only lower a key's RPM, and set its other limits.

The same class covers a different case: a provider rate-limiting us upstream. That variant fails over to another deployment automatically and cools that provider's pool; you only see a provider-originated rate_limit (x-tm-error-origin: upstream) when every candidate was throttled, and its retry-after honors the provider's value capped at 60 seconds.

Organization tokens and concurrency

Beyond RPM, every organization and every key has a token-per-minute limit and a concurrency limit, enforced in the same shared window:

ScopeRPMTokens per minuteConcurrent requests
Organization (all keys together)300 (20 until the first purchase)1,000,0008
Keyits ceiling above1,000,0008

These are the defaults. Set your own on /console/limits or with POST /api/admission-limits; a key's RPM cannot go above its ceiling, and a model-scoped limit can be added on top. Under a contract we can replace these defaults, and the keys' RPM ceiling, for your organization; a limit you set yourself then still applies and can only lower them. Tokens are counted as an estimate at admission — the request's UTF-8 bytes plus a media allowance, plus the output bound — and replaced by observed usage when the request completes. A refusal is the same 429 shape as above with limit_scope org or key and limit_kind tpm or concurrency.

Every admitted response carries x-tm-remaining-rpm and x-tm-remaining-tpm: the smallest remaining amount across the scopes it claimed, at that moment. GET /v1/limits?model=<id> on the gateway returns the effective limits for your key.

Provider capacity is a further scope: each house deployment has a declared pool, and one organization may use at most half of a pool's usable capacity (three quarters on the decision-model pools). Exhausting it is a 429 with x-tm-error-origin: upstream_quota and x-tm-limit-scope: pool. Details, including workspace and principal scopes, on Limits and capacity.

A dedicated endpoint has its own limits, and they replace the defaults above for its traffic. They can include a requests-per-second limit: a count per second, or a paced limit that holds a short burst in a queue. A paced request that waited carries x-tm-queue-ms, the wait in milliseconds. Above the hard rate the refusal is x-tm-limit-kind: rps; when the wait would be longer than its bound, rps_queue. Its refusals carry x-tm-limit-scope: endpoint.

Monthly spend caps

Two optional hard caps, set in the console:

Both run on the UTC calendar month and reset at the first instant of the next month. They overlap freely; the strictest one wins. In the console, leaving a cap field blank changes nothing and setting it to 0 removes the cap. Keys near their cap are flagged, and per-key month-to-date spend is shown against the cap.

How enforcement works — reserve-aware, refuse-early

Before dispatching anything, the gateway reserves a worst-case cost for the request (see the math below). The cap check then asks: would settled spend this month + the worst-case cost of everything currently in flight, this request included exceed the cap? If yes, the request is refused before it costs anything. The in-flight side is claimed synchronously, so concurrent requests cannot race past the cap together.

The consequence to design around: the check is conservative. A request near the boundary can be refused even though its eventual settled cost would have squeezed under — the cap refuses early rather than overshoot. A tight max_tokens shrinks the worst-case reserve and buys back headroom near the cap.

The refusal

A cap breach is the standard out-of-credits 429 (insufficient_quota — the class OpenAI SDKs already recognize natively), with the exact reset instant in both the message and a header:

HTTP/1.1 429 Too Many Requests
x-tm-error-code: insufficient_quota
x-tm-cap-reset: 2026-10-01T00:00:00.000Z
json
{
  "error": {
    "code": "insufficient_quota",
    "type": "insufficient_quota",
    "message": "org monthly cap exceeded; resets 2026-10-01T00:00:00.000Z",
    "metadata": { "error_type": "insufficient_quota", "origin": "gateway_admission", "retryable": true }
  },
  "request_id": "..."
}

The message names the scope that tripped (org or key). On the Anthropic surface the body is a native billing_error carrying the same error_type and message. Do not retry before the reset time — no amount of backoff changes a hard cap, whatever retryable says.

The worst-case reserve

The number the cap and balance checks use is deliberately pessimistic:

InputValue
Estimated input tokensone per UTF-8 byte of the request body, plus 65,536 per image or document part; an image sent inline as base64 counts the 65,536 only, not its bytes
Assumed output tokensyour max_tokens (or max_completion_tokens); 4096 when omitted, and written into the request; more than 32,768 is a 400
Images (/v1/images/generations)assumed output = n × the model's per-image ceiling
Roundingreserve rounds up; settlement rounds down — both in your favor

The reserve exists only while the request is in flight: at settlement it is released and replaced by the observed cost, computed from provider-reported usage at the prices pinned when the request was admitted. Your balance and caps are only ever debited for observed usage — with one exception: an attempt whose outcome the gateway could not observe (unknown_after_crash) settles at the reserved maximum, because the call may have been billed upstream (see Pricing & billing).

Balance exhaustion

Every request must be coverable in the worst case: the org balance has to cover the worst-case reserve of everything the org has in flight, this request included. When it cannot, the refusal is the same insufficient_quota 429 with a different message:

json
{
  "error": {
    "code": "insufficient_quota",
    "type": "insufficient_quota",
    "message": "insufficient marketplace credits for authorized house fallback",
    "metadata": { "error_type": "insufficient_quota", "origin": "gateway_admission", "retryable": true }
  },
  "request_id": "..."
}

To tell the two insufficient_quota cases apart in code: x-tm-cap-reset is present only on cap breaches.

Tip

Signup does not add free credit. Buy credits in Billing; eligible purchases may receive matching promotional credit. If calls return insufficient_quota, check your balance and spend caps there.

Handling the two 429s

Both arrive as 429, so the OpenAI SDK raises RateLimitError for both — switch on error_type, because the correct reactions are opposite:

python
import os, time
from openai import OpenAI, RateLimitError

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])

def create(**kw):
    try:
        return client.chat.completions.create(**kw)
    except RateLimitError as e:
        kind = e.response.json()["error"]["metadata"]["error_type"]
        if kind == "rate_limit":
            time.sleep(int(e.response.headers.get("retry-after", "10")))
            return client.chat.completions.create(**kw)  # one retry
        raise  # insufficient_quota: retrying cannot help — credits or a raised cap
typescript
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY });

try {
  await client.chat.completions.create({
    model: "claude-haiku-4-5",
    messages: [{ role: "user", content: "hello" }],
  });
} catch (err) {
  if (err instanceof OpenAI.APIError && err.status === 429) {
    const kind = (err.error as { metadata?: { error_type?: string } })?.metadata?.error_type;
    if (kind === "rate_limit") {
      // wait err.headers["retry-after"] seconds, retry once
    } else {
      // insufficient_quota — stop; x-tm-cap-reset (cap breaches only) says when a cap reopens
      const capReset = err.headers?.["x-tm-cap-reset"];
    }
  }
}

Other edges

Being straight about the edges, so you can plan around them rather than discover them:

  • No daily or weekly caps. The only spend window is the UTC calendar month. If you need a tighter blast radius today, a low per-key cap on a purpose-made key is the tool.
  • No self-serve RPM raises. Trial keys are 20 RPM, console keys 300 (20 until the organization's first purchase); a limit you set can only lower a key's RPM.
  • Sign-ups that look automated or abusive are refused. The page says that unusual activity was detected; contact support if this happens to you. Existing accounts sign in without limit.
  • Request size: bodies over 10 MB get a 413 request_too_large.
  • Image response size: an upstream image response over 32 MiB is a 502 gateway_error and bills nothing — lower n or the quality. It is not a 413.
  • An image request holds one admission slot for its whole render (typically 10–60 s), because there is no first byte to commit on.

When any of this changes it will change here first — this page is kept in lockstep with the enforcing code.

Markdown source for agents: /docs/limits.md · index at /llms.txt