# Rate limits & spend caps

Several independent brakes can stop a request before it reaches a provider: the key's requests-per-minute limit, the organization's token and concurrency limits, the provider pool's capacity, the monthly spend caps, and the balance itself. Each refuses with a typed error that names its class in the body and in `x-tm-error-code` — nothing is throttled or queued silently, and a refused request costs nothing. The error shapes follow the same conventions as everything else — see [Errors](/docs/errors). Pools, fairness and what happens when the shared store is down are on [Limits and capacity](/docs/admission).

## Per-key requests per minute

Every key has an RPM limit, enforced over a rolling 60-second window per key. The window lives in a shared store, so every gateway instance counts the same requests, and a request's charge can stay in the window for up to 61 seconds.

| Key | RPM |
|---|---|
| First trial key shown after verified browser sign-in | 20 |
| Key created in the console or with the identity API | 300 |
| The console's managed Playground key | 60 |

Until the organization has bought credit (or holds credit we granted), every customer key and the organization as a whole run at the trial rate: 20 requests per minute, however many keys it creates. The first purchase lifts it within seconds; no key changes. Requests served only by your own provider keys (BYOK) and the managed Playground key are not held to it. Historical card-verification bonuses remain promotional credit, not purchases.

An organization on trial credit can also be limited when we detect unusual usage. Such a request gets `429 usage_limited` (with `retry-after`) or `403 usage_restricted`, and the message asks you to contact support, who can review the account. A purchase usually lifts it.

Over the limit, the request is refused before any reservation or provider work:

```
HTTP/1.1 429 Too Many Requests
retry-after: 1
x-tm-error-code: rate_limit
x-tm-error-origin: gateway_admission
x-tm-limit-scope: key
x-tm-limit-kind: rpm
x-tm-limit-id: key:<key id>
```

```json
{
  "error": {
    "code": "rate_limit",
    "message": "configured quota exhausted",
    "type": "invalid_request_error",
    "metadata": {
      "error_type": "rate_limit",
      "origin": "gateway_admission",
      "limit_scope": "key",
      "limit_kind": "rpm",
      "limit_id": "key:<key id>",
      "retryable": true,
      "retry_at": "2026-09-24T10:14:04.000Z"
    }
  },
  "request_id": "..."
}
```

On the Anthropic surface the same refusal is a native `rate_limit_error` carrying the same `error_type` and `metadata`, so the Anthropic SDK's own backoff logic engages.

`retry-after` on a rolling window is advisory: the window frees one second at a time, and other requests may take the room first. Honor it and retry; don't hammer.

> [!NOTE]
> The RPM values are fixed today: there is no self-serve way to raise a key's ceiling. On [`/console/limits`](https://app.routerplus.com/console/limits) you can only lower a key's RPM, and set its other limits.

The same class covers a different case: a provider rate-limiting us upstream. That variant fails over to another deployment automatically and cools that provider's pool; you only see a provider-originated `rate_limit` (`x-tm-error-origin: upstream`) when every candidate was throttled, and its `retry-after` honors the provider's value capped at 60 seconds.

## Organization tokens and concurrency

Beyond RPM, every organization and every key has a token-per-minute limit and a concurrency limit, enforced in the same shared window:

| Scope | RPM | Tokens per minute | Concurrent requests |
|---|---|---|---|
| Organization (all keys together) | 300 (20 until the first purchase) | 1,000,000 | 8 |
| Key | its ceiling above | 1,000,000 | 8 |

These are the defaults. Set your own on [`/console/limits`](https://app.routerplus.com/console/limits) or with `POST /api/admission-limits`; a key's RPM cannot go above its ceiling, and a model-scoped limit can be added on top. Under a contract we can replace these defaults, and the keys' RPM ceiling, for your organization; a limit you set yourself then still applies and can only lower them. Tokens are counted as an estimate at admission — the request's UTF-8 bytes plus a media allowance, plus the output bound — and replaced by observed usage when the request completes. A refusal is the same 429 shape as above with `limit_scope` `org` or `key` and `limit_kind` `tpm` or `concurrency`.

Every admitted response carries `x-tm-remaining-rpm` and `x-tm-remaining-tpm`: the smallest remaining amount across the scopes it claimed, at that moment. `GET /v1/limits?model=<id>` on the gateway returns the effective limits for your key.

Provider capacity is a further scope: each house deployment has a declared pool, and one organization may use at most half of a pool's usable capacity (three quarters on the decision-model pools). Exhausting it is a 429 with `x-tm-error-origin: upstream_quota` and `x-tm-limit-scope: pool`. Details, including workspace and principal scopes, on [Limits and capacity](/docs/admission).

A [dedicated endpoint](/docs/dedicated-endpoints) has its own limits, and they replace the defaults above for its traffic. They can include a requests-per-second limit: a count per second, or a [paced limit](/docs/dedicated-endpoints#the-paced-limit) that holds a short burst in a queue. A paced request that waited carries `x-tm-queue-ms`, the wait in milliseconds. Above the hard rate the refusal is `x-tm-limit-kind: rps`; when the wait would be longer than its bound, `rps_queue`. Its refusals carry `x-tm-limit-scope: endpoint`.

## Monthly spend caps

Two optional hard caps, set in the console:

- an **org-wide** cap covering every key, at [https://app.routerplus.com/console/billing](https://app.routerplus.com/console/billing), and
- a **per-key** cap on any individual key, with **Set cap** at [https://app.routerplus.com/console/keys](https://app.routerplus.com/console/keys).

Both run on the **UTC calendar month** and reset at the first instant of the next month. They overlap freely; the strictest one wins. In the console, leaving a cap field blank changes nothing and setting it to 0 removes the cap. Keys near their cap are flagged, and per-key month-to-date spend is shown against the cap.

### How enforcement works — reserve-aware, refuse-early

Before dispatching anything, the gateway reserves a worst-case cost for the request (see the math below). The cap check then asks: would *settled spend this month + the worst-case cost of everything currently in flight, this request included* exceed the cap? If yes, the request is refused before it costs anything. The in-flight side is claimed synchronously, so concurrent requests cannot race past the cap together.

The consequence to design around: the check is conservative. A request near the boundary can be refused even though its eventual settled cost would have squeezed under — the cap refuses early rather than overshoot. A tight `max_tokens` shrinks the worst-case reserve and buys back headroom near the cap.

### The refusal

A cap breach is the standard out-of-credits 429 (`insufficient_quota` — the class OpenAI SDKs already recognize natively), with the exact reset instant in both the message and a header:

```
HTTP/1.1 429 Too Many Requests
x-tm-error-code: insufficient_quota
x-tm-cap-reset: 2026-10-01T00:00:00.000Z
```

```json
{
  "error": {
    "code": "insufficient_quota",
    "type": "insufficient_quota",
    "message": "org monthly cap exceeded; resets 2026-10-01T00:00:00.000Z",
    "metadata": { "error_type": "insufficient_quota", "origin": "gateway_admission", "retryable": true }
  },
  "request_id": "..."
}
```

The message names the scope that tripped (`org` or `key`). On the Anthropic surface the body is a native `billing_error` carrying the same `error_type` and message. Do not retry before the reset time — no amount of backoff changes a hard cap, whatever `retryable` says.

## The worst-case reserve

The number the cap and balance checks use is deliberately pessimistic:

| Input | Value |
|---|---|
| Estimated input tokens | one per UTF-8 byte of the request body, plus 65,536 per image or document part; an image sent inline as base64 counts the 65,536 only, not its bytes |
| Assumed output tokens | your `max_tokens` (or `max_completion_tokens`); 4096 when omitted, and written into the request; more than 32,768 is a 400 |
| Images (`/v1/images/generations`) | assumed output = `n` × the model's per-image ceiling |
| Rounding | reserve rounds **up**; settlement rounds **down** — both in your favor |

The reserve exists only while the request is in flight: at settlement it is released and replaced by the observed cost, computed from provider-reported usage at the prices pinned when the request was admitted. Your balance and caps are only ever debited for observed usage — with one exception: an attempt whose outcome the gateway could not observe (`unknown_after_crash`) settles at the reserved maximum, because the call may have been billed upstream (see [Pricing & billing](/docs/pricing)).

## Balance exhaustion

Every request must be coverable in the worst case: the org balance has to cover the worst-case reserve of *everything the org has in flight*, this request included. When it cannot, the refusal is the same `insufficient_quota` 429 with a different message:

```json
{
  "error": {
    "code": "insufficient_quota",
    "type": "insufficient_quota",
    "message": "insufficient marketplace credits for authorized house fallback",
    "metadata": { "error_type": "insufficient_quota", "origin": "gateway_admission", "retryable": true }
  },
  "request_id": "..."
}
```

To tell the two `insufficient_quota` cases apart in code: `x-tm-cap-reset` is present **only** on cap breaches.

> [!TIP]
> Signup does not add free credit. Buy credits in [Billing](https://app.routerplus.com/console/billing); eligible purchases may receive matching promotional credit. If calls return `insufficient_quota`, check your balance and spend caps there.

## Handling the two 429s

Both arrive as 429, so the OpenAI SDK raises `RateLimitError` for both — switch on `error_type`, because the correct reactions are opposite:

```python
import os, time
from openai import OpenAI, RateLimitError

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])

def create(**kw):
    try:
        return client.chat.completions.create(**kw)
    except RateLimitError as e:
        kind = e.response.json()["error"]["metadata"]["error_type"]
        if kind == "rate_limit":
            time.sleep(int(e.response.headers.get("retry-after", "10")))
            return client.chat.completions.create(**kw)  # one retry
        raise  # insufficient_quota: retrying cannot help — credits or a raised cap
```

```typescript
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.routerplus.com/v1", apiKey: process.env.TM_API_KEY });

try {
  await client.chat.completions.create({
    model: "claude-haiku-4-5",
    messages: [{ role: "user", content: "hello" }],
  });
} catch (err) {
  if (err instanceof OpenAI.APIError && err.status === 429) {
    const kind = (err.error as { metadata?: { error_type?: string } })?.metadata?.error_type;
    if (kind === "rate_limit") {
      // wait err.headers["retry-after"] seconds, retry once
    } else {
      // insufficient_quota — stop; x-tm-cap-reset (cap breaches only) says when a cap reopens
      const capReset = err.headers?.["x-tm-cap-reset"];
    }
  }
}
```

## Other edges

Being straight about the edges, so you can plan around them rather than discover them:

- **No daily or weekly caps.** The only spend window is the UTC calendar month. If you need a tighter blast radius today, a low per-key cap on a purpose-made key is the tool.
- **No self-serve RPM raises.** Trial keys are 20 RPM, console keys 300 (20 until the organization's first purchase); a limit you set can only lower a key's RPM.
- **Sign-ups that look automated or abusive are refused.** The page says that unusual activity was detected; contact support if this happens to you. Existing accounts sign in without limit.
- **Request size**: bodies over 10 MB get a 413 `request_too_large`.
- **Image response size**: an upstream image response over 32 MiB is a 502 `gateway_error` and bills nothing — lower `n` or the quality. It is not a 413.
- **An image request holds one admission slot for its whole render** (typically 10–60 s), because there is no first byte to commit on.

When any of this changes it will change here first — this page is kept in lockstep with the enforcing code.
