# Agent integration guide

This page is for an AI coding agent that builds RouterPlus into a product's code. It covers the whole integration in order: the key, the SDK, model calls, provider pinning, sessions and prompt caching, errors, limits and cost. Each step links to the reference page that has every detail.

Every docs page is also raw markdown: add `.md` to its URL, for example `https://app.routerplus.com/docs/errors.md`. The index of all pages is `https://app.routerplus.com/llms.txt`. All pages in one file: `https://app.routerplus.com/llms-full.txt`.

> [!NOTE]
> **For agents.** Do not guess model ids, request fields or headers. Copy them from this page or from the live catalog. If an error message disagrees with your plan, trust the error message and read the page it names. Some steps need a person. They are listed in [Steps only a human can do](#steps-only-a-human-can-do). Stop and ask at those steps.

## The short version

1. Point the OpenAI SDK at `https://api.routerplus.com/v1`, or the Anthropic SDK at `https://api.routerplus.com`. Read the key from `TM_API_KEY`, on the server only.
2. Copy model ids exactly from the catalog. The gateway never substitutes a model: an unknown id is a 404.
3. Send `max_tokens` on every request, from 1 to 32,768. Without it, the gateway uses 4,096.
4. When you need provider features such as prompt caching, call Claude models on `/v1/messages` and other models on `/v1/chat/completions`.
5. Keep the start of the prompt identical from turn to turn, and only append. This keeps prompt caches warm.
6. Pin a provider only when you must, with the `provider` object. Check the plan first with `POST /v1/route`.
7. Switch on `error_type`, never on message text. Retry `rate_limit` after `retry-after`. Stop on `insufficient_quota` and tell a human.
8. There is no idempotency key. A retry is a new request, and it is billed again if it reaches a provider.
9. Keep at most 8 requests in flight per organization, unless you raised that limit.
10. Log `x-request-id` and `usage.cost` for every call.

## At a glance

| Item | Value |
|---|---|
| OpenAI-compatible base URL | `https://api.routerplus.com/v1`, for `POST /v1/chat/completions` |
| Anthropic-compatible base URL | `https://api.routerplus.com`, for `POST /v1/messages` (the SDK adds `/v1`) |
| API key | `tm_vk_` followed by 48 hex characters |
| Auth header | `Authorization: Bearer <key>` or `x-api-key: <key>`, on every route |
| Model list | `GET https://api.routerplus.com/v1/models` with your key, or `https://app.routerplus.com/api/models.json` (public, with prices) |
| Cost of a call | `usage.cost`, in USD, in every billed response |
| Audit of a call | the `x-request-id` response header, then `GET https://api.routerplus.com/v1/generation?id=<id>` |
| Images, video and decisions | `POST /v1/images/generations`, `POST /v1/videos` and `POST /v1/decisions` |

### What is not available

Do not build on these. They do not exist on the gateway today:

- The OpenAI Responses API, embeddings, audio, moderation, files and batches. These paths return 404 `not_found`. Use Chat Completions or Messages.
- Server-side conversation memory and idempotency keys. (Session ids exist: see [Sessions and caching](#sessions-and-caching).)
- More than one answer per request (`n` above 1 is a 400).
- Model aliases and model fallback lists. OpenRouter's `models` array and `route` field are dropped.
- Sorting providers by price or speed, and price limits. `provider.sort` and `provider.max_price` are a 400.
- Provider keys or endpoints in a request. `api_key`, `base_url` and `connection_id` in the body are a 400.

## 1. Get a key and keep it safe

| Key | Where it comes from | Requests per minute | Use it for |
|---|---|---|---|
| First trial key | Shown after verified browser sign-in through Clerk | 20 | The first test calls |
| Console key | [https://app.routerplus.com/console/keys](https://app.routerplus.com/console/keys) | 300 | Production, staging and CI |
| Identity API key | `POST https://app.routerplus.com/api/identity/keys` | 300 | One key for each end user or service (step 7) |

To get a key, the user opens [Sign up](https://app.routerplus.com/signup), completes Clerk
authentication and email verification, and saves the first key shown after
sign-in. Add paid credits in [Billing](https://app.routerplus.com/console/billing); signup does not grant free credit. Existing customers sign in
at [https://app.routerplus.com/login](https://app.routerplus.com/login) using the same verified email and
create a key at [API keys](https://app.routerplus.com/console/keys). Each raw key is shown once.

We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match.

Browser signup is required; the former programmatic signup and local magic-link
submission routes have been removed. Once `TM_API_KEY` is available, the agent
can configure clients and make API calls normally.

Rules for the key:

1. Keep the key on the server. Read it from the `TM_API_KEY` environment variable or from your secret store. Never commit it, log it or put it in a URL.
2. Do not put the key in a browser or a mobile app. Anyone can copy it from there. The gateway accepts cross-origin requests, so a copied key works from any web page. If a browser-only internal tool must call the gateway, give it its own key with a low monthly spend cap.
3. Use one key for each service and each environment. Spend, caps and logs are kept per key, and you can disable one key without stopping the others.
4. Use a console key in production. A trial key allows only 20 requests per minute.
5. To rotate a key: create the new key, deploy it, then disable the old key in the console. A disabled key stops working within about 5 seconds. The console cannot enable it again. An owner or admin can, with the identity API.

Details: [Authentication](authentication.md).

## 2. Connect the SDK

Change two values in the client you already use: the base URL and the key. Set both in code. Do not depend on `OPENAI_BASE_URL` or `ANTHROPIC_BASE_URL` in the environment. When that variable is missing, the SDK sends the request, with its key, to the provider's own API instead of the gateway.

Python — the OpenAI SDK:

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.routerplus.com/v1",
    api_key=os.environ["TM_API_KEY"],
    # Optional: these name your app in usage analytics. They carry no content.
    default_headers={"HTTP-Referer": "https://your-app.example", "X-Title": "Your App"},
)
```

TypeScript — the OpenAI SDK:

```typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.routerplus.com/v1",
  apiKey: process.env.TM_API_KEY,
  defaultHeaders: { "HTTP-Referer": "https://your-app.example", "X-Title": "Your App" },
});
```

With the Anthropic SDK, the base URL has no `/v1`:

```python
import os
import anthropic

client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])
```

```typescript
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({ baseURL: "https://api.routerplus.com", apiKey: process.env.TM_API_KEY });
```

### Which surface to call

Every chat model answers on both surfaces, because the gateway translates between the two formats. But a request keeps all its provider features only when your surface matches the format of the route that serves it. Then the gateway forwards the request as you sent it. The table gives the native surface of each model family's first route:

| Model ids | Native surface | Features that need it |
|---|---|---|
| `claude-*` | `/v1/messages` (Anthropic SDK) | Prompt caching with `cache_control`, `thinking`, server tools, image input |
| `gpt-*` | `/v1/chat/completions` (OpenAI SDK) | `response_format`, `reasoning_effort`, `seed`, `logprobs`, image input |
| `author/model` ids, for example `deepseek/deepseek-v4-flash` | `/v1/chat/completions` (OpenAI SDK) | The same OpenAI fields, where the host supports them |

On the other surface, the gateway translates. A field that the translation cannot carry is dropped and named in the `x-tm-dropped-params` response header. If dropping the field would change the answer, the request is a 400 that names the field instead. Before it refuses, the gateway tries another provider of the same model that speaks your format. The `x-tm-provider` header names the provider that served you. The catalog gives each model's native format in `providers[0].dialect`. If your code sends only text and function tools, either surface is fine for every model.

The native surface helps only while the first route serves you. A Claude request on `/v1/messages` that is pinned to `openrouter`, or fails over to it, is translated: `thinking`, server tools and image blocks cannot cross (the request is a 400 if no other route can take it), and other Anthropic-only fields such as `top_k`, `metadata` and `cache_control` are dropped and named.

Details: [Wire compatibility](compat.md).

## 3. Call models

### Rules for every request

1. **Model ids are exact.** Copy them from `GET /v1/models` or `/api/models.json`. Claude and GPT ids are bare: `claude-sonnet-5`, `gpt-4o-mini`. Open models keep their author prefix: `moonshotai/kimi-k3`. OpenRouter variants such as `:free` do not exist here. A dated snapshot of a listed model, such as `claude-haiku-4-5-20251001`, also works: it routes and bills as its model.
2. **Keep model ids in configuration**, in one place. Check them against `GET /v1/models` when your service starts. Then a model change is a configuration change.
3. **Use the right route for the output.** Chat models answer on the two chat surfaces. Image models answer only on `POST /v1/images/generations`, video models only on `POST /v1/videos`, and decision models only on `POST /v1/decisions`. The catalog field `output_modalities` (`text`, `image`, `video` or `decisions`) tells you which is which. A decision model also takes its own question format: the catalog field `decisions_wire` in `/api/models.json` names it (`systemone` for Jev, Mercury Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider v1 27B and Sage, and `gliner`), and [POST /v1/decisions](/docs/api-decisions) gives each one. Do not guess the format from the model's name.
4. **Send `max_tokens` every time**, from 1 to 32,768. If you leave it out, the gateway writes `max_tokens: 4096` into the request, and a longer answer stops with `finish_reason: "length"`. A value above 32,768 is a 400. The gateway also holds balance for `max_tokens` while the call runs, so a realistic value helps on a small balance.
5. **Ask for one answer.** `n` above 1 is a 400. Make separate calls for several samples.
6. **Send the whole conversation on every turn.** The gateway keeps no conversation state, and it stores no prompts or answers.
7. **Keep `temperature` at 1 or less** in code that can reach Claude models. On a translated request, a higher value is a 400.
8. **Refuse drops where they matter.** A field outside the forwarded set is dropped and named in `x-tm-dropped-params`. Send `"provider": {"require_parameters": true}` to get a 400 instead of a drop. The 400 comes from the first route that would drop a field: the gateway does not go on to a later route that could carry it. So combine it with `only` or `order`. For example, the `anthropic` route drops `seed`, so for `seed` on a Claude model over `/v1/chat/completions`, send `{"only": ["openrouter"], "require_parameters": true}`.

### Streaming

Stream answers that a person watches, and long answers. The OpenAI surface sends, in this order: a role chunk, content chunks, a finish chunk, a usage chunk with an empty `choices` list and `usage.cost`, then `data: [DONE]`. The Anthropic surface sends the native event sequence, with usage and `cost` in the final `message_delta`.

- If the provider fails after the answer starts, the stream ends with one error event and nothing after it (no `[DONE]`). The SDKs raise an error. The partial answer is billed. The gateway never continues an answer on another provider. To try again, send the whole turn again.
- During a silence of 15 seconds or more, the gateway sends a keep-alive: an SSE comment on the OpenAI surface, a `ping` event on the Anthropic surface. The SDKs skip them. If you parse SSE yourself, skip them too.
- The first response headers must arrive within 20 seconds, or the gateway tries the next provider. A request can run for at most 450 seconds.
- To stop an answer, close the connection. The provider stops. You pay for the usage that the provider reported. If it reported none, you pay for an estimate of the text already streamed, at about four characters per token. If no text streamed yet, the gateway sees no usage, and you pay the full hold for the call: the input estimate plus `max_tokens`, at the model's prices.

Python — the OpenAI SDK, keeping the cost and handling a failure in the middle of an answer:

```python
from openai import APIError

parts, usage = [], None
try:
    stream = client.chat.completions.create(
        model="claude-sonnet-5",
        max_tokens=2048,
        stream=True,
        messages=messages,
    )
    for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content:
            parts.append(chunk.choices[0].delta.content)
        if chunk.usage:  # the usage chunk; when usage is known, it also comes before an error
            usage = chunk.usage
except APIError:
    # The answer failed part way. What arrived is billed. Do not join a retry onto
    # `parts`: send the whole turn again, or show the error.
    raise
print("".join(parts), "cost:", usage.cost if usage else None)
```

Details: [Streaming](streaming.md).

### Calls that do not stream

Set the client timeout above 450 seconds. The SDK default of 10 minutes is fine. The Anthropic SDKs refuse a call that does not stream when `max_tokens` is above about 21,000, unless you pass an explicit `timeout`. Stream those calls instead.

Do not close the connection of a call that does not stream. The gateway stops the provider, but it sees no usage, so you pay the full hold for the call: the input estimate plus `max_tokens`, at the model's prices. If you may need to cancel, stream the call. An image render is different: it does not stop. It finishes and is billed at the usage that the provider reports.

### Tools, reasoning and structured output

- Function tools work on both surfaces for every chat model. The gateway translates tool calls and tool results.
- Anthropic server tools (web search, computer use, bash, text editor) work only for Claude models on `/v1/messages`, while the `anthropic` route serves them. The `anthropic` route refuses `mcp_servers` and `container`. The request then falls over to `openrouter`, which runs it without them and names them in `x-tm-dropped-params`. Send `"provider": {"require_parameters": true}` to get a 400 instead.
- `thinking` works only for Claude models on `/v1/messages`, while the `anthropic` route serves them. Reasoning comes back in different fields. On `/v1/chat/completions`, Claude's thinking arrives as `reasoning_content`, and open models return OpenRouter's own `reasoning` field (and `reasoning_details`). On `/v1/messages`, a streamed answer from a non-Claude model carries its reasoning as `thinking` blocks, and a non-streamed answer does not carry it. You can send `thinking` blocks back in the next turn.
- `response_format` needs an OpenAI-format provider: GPT and `author/model` ids. For JSON from a Claude model, force a tool call with `tool_choice` and read the tool input.
- Do not end the message list with an assistant message when a Claude model can serve the request from the OpenAI surface. That is a 400, because Claude rejects a prefilled answer.

### Images in prompts

Send image parts on the model's native surface: Claude on `/v1/messages`, GPT and `author/model` ids on `/v1/chat/completions`. When the gateway must translate, an image part is a 400 that names the part. Send an image, document, audio or video part only to a model that takes it: `architecture.input_modalities` in `GET /v1/models` lists each model's inputs. A part the model does not take is a 400 that names the part, and nothing is billed. Each image or document part counts as 65,536 tokens against your limits and your balance hold while the call runs.

### Other routes

- **Images:** `POST /v1/images/generations` takes the OpenAI Images API body. The image comes back as base64 in the response (`response_format: "url"` is a 400). `n` is 1 unless the model allows more: see `supported_parameters.n` in `/api/models.json` (GPT Image models allow up to 4). There is no streaming. A failed render bills nothing. See [POST /v1/images/generations](api-images.md).
- **Video:** `POST /v1/videos` starts a job. Poll `GET /v1/videos/{id}` until it ends, then download `GET /v1/videos/{id}/content`. A failed job costs nothing. See [POST /v1/videos](api-videos.md).
- **Decisions:** `POST /v1/decisions` sends a state and typed questions to a decision model: `typesafe/jev-1.13` (Jev, from TypeSafe), `inception/mercury-decide` (Mercury Decide, from Inception), `bespokelabs/nimble-v3` (Bespoke Nimble v3, from Bespoke Labs), `cloudflare/clef` and `cloudflare/clef-flash` (Clef and Clef-flash, from Cloudflare), `routerplus/decider-2b` and `routerplus/kev-4b` (Decider 2B and Kev 4B, from RouterPlus), `perplexity/pplx-decider-v1-27b` (Perplexity Decider v1 27B, from Perplexity), `levanto/sage-1.2` (Sage, from Levanto) or `fastino/gliner-2.5-decide` (GLiNER-2.5-Decide, from Fastino). It returns one typed answer for each question, such as a choice, a score or a yes/no probability. The question format follows the model: System One questions (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage) have a `type`, GLiNER takes a `schema` of classifications and extractions in place of `questions`, and a question in another model's format is a 400. GLiNER's numbers are confidences, not calibrated probabilities. Perplexity Decider bills the state once per question, so on it ask only the questions you need. The gateway never fails over from one model to another. Use it for routing, classification and moderation points in your code, in place of a chat model that you parse. See [POST /v1/decisions](api-decisions.md).
- **Token counting:** `POST /v1/messages/count_tokens` is free, for models that have an Anthropic-format provider. It still counts against your requests per minute.
- **Your deployed models:** an id such as `tm/<name>-v1` from [Optimize](model-search.md) works on `/v1/chat/completions` only, with keys of the organization that deployed it. With `stream: true` it returns the finished answer as one chunk.

## 4. Pin providers (only when you must)

How routing works without a pin:

- Each chat model has a fixed, ordered list of providers. Today Claude models go to `anthropic` first and `openrouter` second. GPT models go to `openai` first and `openrouter` second. Open models (`author/model` ids) go to `openrouter` only.
- The order does not change from request to request. There is no random spread across providers.
- The gateway moves to the next provider only when the current one fails before it sends any output, or has no capacity left.
- OpenRouter picks the host that runs an open model, for example Fireworks, Together or DeepInfra, for each request.
- The price of a model is the same whichever provider serves it.

Pin when a rule requires one provider, when a feature exists at one provider only, when a session must stay on one host (step 5), or when you compare providers in a test. A pin can cost you fallbacks: `only` and `"allow_fallbacks": false` remove the fallback provider, and `upstream` removes the fallback hosts inside OpenRouter. `order` keeps every fallback.

### The `provider` object

Send it in the request body on `/v1/chat/completions` or `/v1/messages`. The gateway reads it and never forwards it. An unknown key inside it is a 400.

| Field | Value | Effect |
|---|---|---|
| `order` | provider ids | Try these providers first, in this order. The others stay as fallbacks. |
| `only` | provider ids | Use only these providers. |
| `ignore` | provider ids | Never use these providers. |
| `allow_fallbacks` | boolean | `false` keeps only the first route. If that route is resting after failures, the request fails with 503 and `retry-after` instead of moving on. |
| `upstream` | host tags, up to 8 | On an OpenRouter route: use only these hosts, in this order, and never another host. If they all fail, the request fails. |
| `require_parameters` | boolean | `true`: when the first route that would serve you would drop one of your request fields, the whole request is a 400. The gateway does not try a later route. |
| `connections`, `connection_order`, `funding` | ids, `"byok"` or `"house"` | Choose among your own provider connections and who pays. See [Routing policies](routing-policies.md). |
| `regions`, `zdr`, `data_collection` | region list, boolean, `"allow"` or `"deny"` | Hard filters. A route passes only with operator evidence, and marketplace routes have none today. So on marketplace traffic, `regions`, `"zdr": true` or `"data_collection": "deny"` removes every route. |

`order`, `only` and `ignore` take route labels and OpenRouter host tags. The marketplace routes are `anthropic`, `openai` and `openrouter`: the `providers[].id` values in `/api/models.json`. Your own connections use their profile: `openai`, `anthropic`, `azure` or `bedrock`.

- A label that matches a route label selects or orders that route, even when that route does not serve the model. Route labels win over host tags with the same spelling: `only: ["anthropic"]` means the Anthropic route, so on an open model it leaves no route.
- Any other label is a host tag. When the `openrouter` route runs, the gateway sends the tags to OpenRouter as its own `order`, `only` and `ignore`, with your `allow_fallbacks` (default true). OpenRouter then prefers, limits or skips those hosts, and it can still fall back to other hosts unless `allow_fallbacks` is `false`.
- A host tag in `only` keeps the `openrouter` route open and closes the routes that `only` does not name.
- `upstream` is the strict pin: those hosts only, in that order. Host tags from `order`, `only` and `ignore` do not go with it.
- `POST /v1/route` shows the host preferences in the field `openrouter_provider` (`null` when there are none).

Host tags are OpenRouter's provider slugs, for example `fireworks`, `together`, `deepinfra`, `amazon-bedrock` or `google-vertex`. The hosts that run a model are its `served_by` entries in `/api/models.json`. Use only those hosts: the gateway tells OpenRouter to skip some hosts that OpenRouter lists, and a request that names only those hosts fails. For most hosts, `served_by[].id` is the same as the tag. Three differ: `amazon` (tag `amazon-bedrock`), `moonshot` (tag `moonshotai`) and `zai` (tag `z-ai`).

| Goal | `provider` |
|---|---|
| Claude only from Anthropic's own API | `{"only": ["anthropic"]}` |
| Claude through OpenRouter first, Anthropic as the fallback | `{"order": ["openrouter"]}` |
| Claude on Amazon Bedrock | `{"only": ["openrouter"], "upstream": ["amazon-bedrock"]}` |
| An open model on one host for every turn | `{"upstream": ["fireworks"]}` |
| An open model on one host, with one named backup host | `{"upstream": ["fireworks", "together"]}` |
| Fail instead of falling back to another provider | add `"allow_fallbacks": false` |

Check a pin before you ship it. `POST /v1/route` returns the ordered candidates and each excluded route with its reason. It calls no provider and costs nothing. If your controls exclude every route, a real request returns 404 `model_unavailable`, the same error as an unknown model. So when a pinned request gets a 404, run `/v1/route`.

```bash
curl -s https://api.routerplus.com/v1/route \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5","provider":{"only":["openrouter"],"upstream":["amazon-bedrock"]}}'
```

Python — the OpenAI SDK sends fields it does not know through `extra_body`:

```python
reply = client.chat.completions.create(
    model="moonshotai/kimi-k3",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}],
    extra_body={"provider": {"upstream": ["fireworks"]}},
)
```

TypeScript — pass a variable, so the field the SDK types do not know is kept:

```typescript
const request = {
  model: "moonshotai/kimi-k3",
  max_tokens: 1024,
  messages: [{ role: "user" as const, content: "Hello" }],
  provider: { upstream: ["fireworks"] }, // read by the gateway, never forwarded
};
const reply = await client.chat.completions.create(request);
```

Check each answer with its response headers: `x-tm-provider` names the provider, `x-tm-served-by` names the host on an OpenRouter route as OpenRouter spells it (for example `Amazon Bedrock`, not the tag `amazon-bedrock`), `x-tm-upstream-model` gives the id the provider received when it differs from yours, and `x-tm-attempts` counts provider attempts. `GET /v1/generation?id=<x-request-id>` shows the route of every attempt.

To pin for a whole organization, workspace or key instead of each request, save a routing policy with the policy API. It takes a console session, so a human sets it up. A request's `provider` object can narrow a saved policy, but it can never add routes. See [Routing policies](routing-policies.md).

### Your own provider keys (BYOK)

You can route through your own OpenAI, Anthropic or Azure key, or an AWS Bedrock role. A human adds it once in the console as a connection. Your code still calls the gateway with a marketplace key. The provider bills you directly, and `usage.cost` is 0. A key bound to a connection can call only that connection's exact model ids, and it does not fall back to marketplace supply unless a routing policy allows it. See [Bring your own key](byok.md).

## 5. Keep sessions on one route and caches warm

A prompt cache saves money and time only when the next request reaches the same provider, or the same host, with the same prompt start. This is how the gateway treats sessions today:

- The gateway keeps no session state. Your app stores the conversation and sends all of it on every turn.
- Routing is deterministic. For the same model, key and `provider` object, every request tries the same route first. So the turns of a session stay on one provider without a session id.
- A turn moves to the next route when the first route fails before it sends output (one failure is enough), when the route's capacity pool is full, or while its provider's 429 pause lasts (the provider's `retry-after`, 1 to 60 seconds). After two failures in a row, later turns skip the route for 30 seconds. A turn that moves can miss the cache once. On a follow-up turn, the gateway first waits up to 5 seconds for a full pool: see [Waiting for the first route](#waiting-for-the-first-route).
- Send a session id, and the gateway gives it to each provider in the provider's own field: see [Sessions and caching](#sessions-and-caching). Claude Code's `x-claude-code-session-id` header counts, with no change on your side.
- For an open model, OpenRouter picks the host. With a session id (or the id that the gateway makes when you send none), OpenRouter keeps the conversation on one host from its first request, but it does not promise to. To force one host, pin it with `upstream`, as shown below. Remember that `upstream` turns off host fallback: if the pinned hosts fail, the request fails.

Rules for a high cache hit rate:

1. Keep the start of the prompt byte-identical: the same system prompt, the same tool definitions in the same order, and earlier messages unchanged. Only append.
2. Put content that changes (the time, ids, documents retrieved for this turn) after the stable part, at the end.
3. Keep the model and the `provider` object the same for the whole session.
4. For Claude, put `cache_control` on the block that ends the stable part (a system, tool or message block), or send the top-level `cache_control` field for automatic caching. Both surfaces keep the marks for Anthropic and OpenRouter: see [Cache marks](#cache-marks). On marketplace routes the cache lasts 5 minutes: `ttl: "1h"` is removed.
5. OpenAI caches long repeated prompt starts by itself. Many open-model hosts do too. You set nothing for them.
6. For an open model, pin one host for each session when cache hits matter more than host fallback.
7. Measure. Cache reads are `usage.prompt_tokens_details.cached_tokens` on the OpenAI surface and `usage.cache_read_input_tokens` on the Anthropic surface. Cache writes are `cache_write_tokens` and `cache_creation_input_tokens`. Cache reads are billed at the lower `cached_prompt` price. Claude cache writes are billed at the higher `cache_write` price. Prices are in `/api/models.json`.

Python — the Anthropic SDK, with a cache mark on a long system prompt that does not change:

```python
reply = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    system=[{
        "type": "text",
        "text": STABLE_INSTRUCTIONS,              # the same bytes on every turn
        "cache_control": {"type": "ephemeral"},   # cache everything up to here
    }],
    messages=history,                             # append only
)
print(reply.usage)  # cache_read_input_tokens, cache_creation_input_tokens and cost
```

The provider caches a prompt start only above a minimum length, so a short system prompt shows no cache reads.

For open models, give each session its own host. The same session always gets the same first host, and different sessions spread across hosts. The second host is used only when the first one fails:

```python
import hashlib

# Hosts that run the model: its served_by entries in https://app.routerplus.com/api/models.json (tags as in step 4)
HOSTS = ["fireworks", "together", "deepinfra"]

def session_route(session_id: str) -> dict:
    i = int(hashlib.sha256(session_id.encode()).hexdigest(), 16) % len(HOSTS)
    return {"upstream": [HOSTS[i], HOSTS[(i + 1) % len(HOSTS)]]}

reply = client.chat.completions.create(
    model="moonshotai/kimi-k3",
    max_tokens=2048,
    messages=history,
    extra_body={"provider": session_route(session_id)},
)
print(reply.usage.prompt_tokens_details.cached_tokens)
```

### Sessions and caching

#### Session ids

On `POST /v1/chat/completions` and `POST /v1/messages`, the session id is the first valid value among these inputs, in this order. Token counting, images and videos do not use it.

| Order | Input | Kind |
|---|---|---|
| 1 | `session_id` | Top-level body field |
| 2 | `x-session-id` | Header |
| 3 | `x-session-affinity` | Header |
| 4 | `x-claude-code-session-id` | Header that Claude Code sends by itself |

- A valid value has 1 to 256 printable ASCII characters. `prompt_cache_key`, a top-level body field, follows the same rule.
- The gateway ignores an invalid value and uses the next input. An invalid value alone is never an error. The ignored input is named in `x-tm-dropped-params` as `session_id (invalid)`, `prompt_cache_key (invalid)`, `header:x-session-id (invalid)`, `header:x-session-affinity (invalid)` or `header:x-claude-code-session-id (invalid)`. With `"require_parameters": true`, that makes the request a 400, as any dropped field does.
- Valid session inputs are never named in `x-tm-dropped-params`, and they never trigger `require_parameters`.
- `prompt_cache_key` stays OpenAI's own field. It wins over the session id on OpenAI routes, and it is the fallback on OpenRouter.
- When two valid inputs disagree, the order decides. There is no error.
- Codex's `session-id` header is not read. It can matter only after the gateway adds `POST /v1/responses`, because current Codex calls only the Responses API.

What each route receives:

| Route | Field | Value |
|---|---|---|
| `openai` (marketplace), and `openai` and `azure` (your connections) | `prompt_cache_key` | Your `prompt_cache_key`, or else the session id. Nothing if you sent neither |
| `openrouter` (marketplace) | `session_id` | The session id, or else your `prompt_cache_key`, or else an id that the gateway makes. If you sent `prompt_cache_key`, it goes too |
| `anthropic` and `bedrock` | Nothing | Anthropic has no session input. Its cache matches the prompt start |

- Values go to the provider unchanged. A value longer than the provider accepts is replaced by `tm-` followed by 43 base64url characters: a hash of the value.
- For OpenRouter only, when you send neither a session id nor `prompt_cache_key`, the gateway makes an id from your organization id, the model and the opening messages, up to the first user message. OpenRouter then keeps the conversation on one host from its first request.
- The gateway keeps no session table. The id goes to the provider, and the provider routes on it.
- `prompt_cache_retention` goes unchanged to `openai` and `azure` routes, both `"in_memory"` and `"24h"`. On other routes it is dropped and named.

#### End-user ids on marketplace routes

On marketplace routes, the provider gets a hash in place of your end user's id, and the raw values do not go. The id is the first string among the body fields `safety_identifier`, `user` and `metadata.user_id`. The hash is `tm-` followed by 43 base64url characters, made from your organization id and that id. Without an id, it is made from your organization id alone. Each marketplace route gets exactly one field:

| Marketplace route | Field that carries the hash | Fields removed |
|---|---|---|
| `openai`, `azure` | `safety_identifier` | `user` |
| `openrouter` | `user` | `safety_identifier` |
| `anthropic` | `metadata.user_id` | `user`, `safety_identifier` |
| every other route (OpenAI-compatible), dedicated capacity included | `user` | `safety_identifier`, `metadata.user_id` |

- A provider that blocks an id then blocks one user of one customer, not the whole marketplace account.
- On your own connections, `user` and `metadata` pass unchanged. `safety_identifier` passes unchanged to `openai` and `azure` connections, and it is dropped and named on `anthropic` and `bedrock` connections.
- Identity fields that a route uses are never named in `x-tm-dropped-params`.

#### Cache marks

The gateway never adds a cache mark that you did not send. Block and part marks (`cache_control`) are handled this way:

| Request | Route | What happens to the marks |
|---|---|---|
| `/v1/chat/completions` | Anthropic direct or Bedrock | Each text part keeps its mark. A marked system part turns the system prompt into blocks. A tool message's last mark goes on its `tool_result` block. Other part keys are not forwarded |
| `/v1/chat/completions` | OpenRouter | Parts pass unchanged |
| `/v1/chat/completions` | OpenAI direct or Azure | Part marks pass unchanged. OpenAI and Azure ignore them |
| `/v1/messages` | Anthropic direct or Bedrock | Block marks pass unchanged |
| `/v1/messages` | OpenRouter | Marks survive in system, user, assistant and tool content. Marks on tool definitions and on `tool_use` blocks are removed and named as `tools[i].cache_control` or `messages[i].content[j].cache_control` |
| `/v1/messages` | OpenAI direct or Azure | Marks are removed and named |

A top-level `cache_control` field (automatic caching), on either surface:

- Anthropic direct and OpenRouter: sent unchanged.
- Bedrock: turned into a mark on the last block that can carry one, counted back from the end of the messages. If no block can carry it, it is dropped and named as `cache_control`.
- OpenAI direct and Azure: dropped and named as `cache_control`.
- Any route: a field that Anthropic would refuse is dropped and named as `cache_control`. Anthropic refuses a field other than `{"type": "ephemeral"}` with an optional `ttl` of `"5m"` or `"1h"`, a fifth mark (four marks exist already), a 1-hour field after a 5-minute mark, and a field whose TTL differs from the mark on the last block.

On marketplace routes, `ttl: "1h"` is removed from every mark, so the cache lasts 5 minutes. The drop is named as `<path>.cache_control.ttl`, for example `system[0].cache_control.ttl`, `messages[2].content[0].cache_control.ttl`, or `cache_control.ttl` for the top-level field. `ttl: "5m"` passes. Your own connections keep `ttl: "1h"`.

#### Waiting for the first route

On a follow-up turn (an assistant or tool message after a user message), if the first route's capacity pool is full, the gateway waits for that pool for up to 5 seconds (an operator setting) before it moves the request to the next route. It never waits past the request's deadline. The `x-tm-affinity-wait-ms` response header gives the wait in milliseconds, only when the wait was more than 0.

## 6. Handle errors and retries

Every failure carries one stable class, `error_type`. Read it from `error.metadata.error_type` on the OpenAI surface, or `error.error_type` on the Anthropic surface. The `x-tm-error-code` header carries it too, but browsers cannot read that header. Never parse the message text. Log the `x-request-id` header with every failure.

| `error_type` | HTTP | Retry | What to do |
|---|---|---|---|
| `rate_limit` | 429 | Yes, once | Wait `retry-after` seconds. If `x-tm-limit-kind` is `concurrency`, send fewer requests at once. |
| `email_not_verified` | 403 | No | Stop and tell a human to open the verify link in their email; the key starts working within seconds of the click. |
| `insufficient_quota` | 429 | No | Stop and tell a human: add credits or raise a spend cap. On a cap, `x-tm-cap-reset` gives the reset time. |
| `upstream_error` | 5xx | Yes, once, after about 2 seconds | The gateway already tried every other provider it could. If the retry fails too, report it with the `x-request-id`. |
| `upstream_error` | 4xx | No | The provider refused the request itself. Read the message. |
| `gateway_error` | 503 | Yes | Wait `retry-after`. The request did not reach a provider. |
| `gateway_error` | 504 | Yes, once | The deadline passed while the gateway tried providers. An earlier attempt may be billed: check `GET /v1/generation?id=<x-request-id>` first. |
| `gateway_error` | 400 | No | Send one output limit, from 1 to 32,768. |
| `gateway_error` | 500 | No | Report it with the `x-request-id`. |
| `model_unavailable` | 404 | No | Fix the model id, or fix your `provider` object (check it with `POST /v1/route`). |
| `model_unavailable` | 502 | Yes, once | The connection to every provider failed. |
| `invalid_request` | 400 | No | Fix the field that the message names. |
| `context_overflow` | 400 | No | Shorten the input, or use a model with a larger context. |
| `content_policy` | 400 | No | Change the prompt. It is never sent to another provider. |
| `auth` | 401 | No | Check the header, the key, and whether the key was disabled. A 401, 402 or 403 relayed from every provider is the marketplace's own account problem: report it with the `x-request-id`. |
| `request_too_large` | 413 | No | Keep the body under 10 MB (1 MB for a video request). |
| `not_found` | 404 | No | Use a route that exists. The list is in [Errors](errors.md). |

Rules:

1. Let the SDK retry. The official OpenAI and Anthropic SDKs retry 429 and 5xx responses twice by default, and they wait for `retry-after`. Keep that default. Do not add a second retry loop around it. The SDKs also retry an `insufficient_quota` response. That costs nothing, but your code must then stop.
2. There is no idempotency key. A retried request is a new request. If it reaches a provider, it is billed again. A refusal by the gateway's own limits (`x-tm-error-origin: gateway_admission`) cost nothing.
3. Never retry `insufficient_quota`, `invalid_request`, `context_overflow`, `content_policy` or `auth` in a loop. The request, or the account, must change first.
4. An error in the middle of a stream is final for that answer. Send the whole turn again. Never join two partial answers.

Python — the OpenAI SDK, reading the class:

```python
from openai import APIStatusError

try:
    reply = client.chat.completions.create(model=MODEL, max_tokens=1024, messages=messages)
except APIStatusError as e:
    error = e.response.json().get("error") or {}
    kind = (error.get("metadata") or {}).get("error_type") or e.response.headers.get("x-tm-error-code")
    request_id = e.response.headers.get("x-request-id")
    if kind == "insufficient_quota":
        notify_operator(request_id, error.get("message"))  # money or a cap: a retry cannot help
    raise
```

Details: [Errors](errors.md).

## 7. Limits, spend caps and your users

Default limits:

| Scope | Requests per minute | Tokens per minute | In flight at once |
|---|---|---|---|
| Trial key | 20 | 1,000,000 | 8 |
| Console or identity API key | 300 | 1,000,000 | 8 |
| Organization, all keys together | 300 | 1,000,000 | 8 |

Rules:

1. Limit how many requests your code sends at once. By default the organization allows 8 in flight. A ninth gets 429 `rate_limit` with `x-tm-limit-kind: concurrency`. Use a semaphore or a fixed pool of workers. To run more at once, a human first raises the organization and key limits in [/console/limits](https://app.routerplus.com/console/limits). A key's requests per minute cannot go above its ceiling (300 for a console key). Each marketplace provider pool also limits one organization to half of the pool's capacity. That refusal carries `x-tm-limit-scope: pool`.
2. Read `x-tm-remaining-rpm` and `x-tm-remaining-tpm` on each response, and slow down before they reach zero. `GET https://api.routerplus.com/v1/limits?model=<id>` returns the limits that apply to your key.
3. While a call runs, token limits count an estimate: one token for each four bytes of the request, rounded up, 65,536 for each image or document part, plus `max_tokens`. An image sent inline as base64 counts its 65,536 only, not its bytes. The real count replaces the estimate when the call ends. Small requests and a realistic `max_tokens` let more calls run at once.
4. Spend caps are monthly, for each key and for the organization, on the UTC calendar month. A human sets them in the console. A cap refuses with 429 `insufficient_quota` and the `x-tm-cap-reset` header. Give each feature or customer its own key and cap, so a bug or a loop cannot spend everything.

Your users:

- By default, one backend key serves all your users. The gateway does not know who your users are, so enforce per-user quotas in your app.
- `safety_identifier`, `user` and `metadata.user_id` identify your end user to the provider. On marketplace routes the gateway sends only a hash of that id with your organization; on your own connections they go unchanged (see [End-user ids on marketplace routes](#end-user-ids-on-marketplace-routes)). They do not identify, limit or bill anyone on the gateway. Do not put emails or names in them.
- For limits that the gateway enforces per user, use the identity API. Create an `end_user` principal for each user, issue a key for it, and set limits on the principal. Your backend keeps those keys. Never give a key to the user. The identity API takes a console session, not an API key, so a human with an owner or admin role sets it up. See [Workspaces and principals](identity.md).
- To name your app in analytics, send the `HTTP-Referer` and `X-Title` headers. They carry no content.

Details: [Rate limits and spend caps](limits.md) and [Limits and capacity](admission.md).

## 8. Track cost

- Every billed response has `usage.cost` in USD. It is the full charge for the request, including provider attempts that failed before the answer. Store it with the `x-request-id`, the model and your own ids, such as the user and the feature.
- In a stream, `cost` is in the usage chunk on the OpenAI surface, and in the last `message_delta` on the Anthropic surface.
- An Anthropic-surface stream that fails in the middle of the answer has no usage in it. Read its cost from `GET /v1/generation?id=<x-request-id>`.
- `GET /v1/usage` returns the balance and the last 20 attempts. For a key bound to its own workspace or principal, the balance fields are `null`. The console shows usage and logs for the whole organization.
- BYOK calls show a cost of 0, because your provider bills you directly.
- To check a charge, multiply the token counts by the prices in `/api/models.json`. The gateway rounds the final charge down.

Details: [Pricing and billing](pricing.md) and [GET /v1/usage and /v1/generation](api-usage.md).

## 9. Verify the integration

One cheap call shows the answer, the cost and the headers that the gateway adds:

```bash
curl -sS -i https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"gpt-4o-mini","max_tokens":16,"messages":[{"role":"user","content":"Reply with OK"}]}'
```

Then confirm that the ledger saw it:

```bash
curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usage
```

The integration is done when all of these are true:

1. Every client uses the gateway base URL and reads the key from `TM_API_KEY` on the server.
2. The key is not in the repository, in a client bundle or in a log.
3. Every model id in the configuration appears in `GET /v1/models`.
4. Every request sends `max_tokens`.
5. One streamed and one non-streamed test call return `usage.cost`.
6. The `x-tm-dropped-params` header is absent, or it names only fields that you accept losing.
7. `x-tm-provider`, and `x-tm-served-by` when you pin a host, show the route that you expect. (`x-tm-served-by` uses OpenRouter's display name, not the tag.)
8. The error handling switches on `error_type`. It stops on `insufficient_quota`, and it adds no retry loop on top of the SDK's own retries.
9. The number of requests in flight stays within the organization's limit.
10. The logs keep `x-request-id` and `usage.cost` for every call.
11. The test calls appear in `GET /v1/usage`.

## Leave instructions for the next agent

Add this block to the `AGENTS.md` or `CLAUDE.md` file of the repository that you integrated. Later coding sessions then keep the integration correct:

```markdown
## LLM calls: RouterPlus gateway

- Every LLM call goes through the RouterPlus gateway. Full guide: https://app.routerplus.com/docs/agent-integration.md
- OpenAI SDK base URL: https://api.routerplus.com/v1. Anthropic SDK base URL: https://api.routerplus.com. Key: the TM_API_KEY environment variable, on the server only.
- Model ids live in configuration. Each one must exist in https://app.routerplus.com/api/models.json. Never invent or alias an id.
- Send max_tokens on every request (1 to 32768). Never send n above 1.
- Claude models: /v1/messages, with cache_control on the stable start of the prompt. GPT and author/model ids: /v1/chat/completions.
- Keep the start of the prompt byte-identical across turns. Only append.
- Send one id per conversation as the x-session-id header (or session_id in the body). Providers then keep its cache warm.
- Provider pins go in the request body "provider" object (only, ignore, order, allow_fallbacks, upstream, require_parameters). Check a pin with POST https://api.routerplus.com/v1/route.
- Errors: switch on error_type. rate_limit: wait retry-after, then retry once. insufficient_quota: stop and tell a human. Never retry invalid_request, context_overflow or content_policy.
- No idempotency key: a retry is a new, billed request. Add no retry loop on top of the SDK's own retries.
- At most 8 requests in flight per organization, unless the limit was raised.
- Log x-request-id and usage.cost for every call.
- Not available: the Responses API, embeddings, audio, files and batches.
```

## Steps only a human can do

- Complete browser signup through Clerk, including email verification, then buy credits in Billing. Signup and card verification alone do not add credit.
- Fix billing when `insufficient_quota` appears: add credits, or raise a spend cap in the console.
- Sign in to the console to create keys, set spend caps and change limits.
- Add your own provider keys (BYOK), save routing policies, and set up workspaces and end-user principals.
- Approve large test runs. Every call is billed.

## Where next

- [Quickstart](/docs/quickstart) — the first streamed call, step by step
- [Errors](/docs/errors) — every error class, and what to do about it
- [Routing policies](/docs/routing-policies) — the provider object, saved policies and BYOK routes
- [Streaming](/docs/streaming) — frame order, keep-alives and failures in the middle of an answer
- [Rate limits & spend caps](/docs/limits) — the numbers behind every 429
- [Wire compatibility](/docs/compat) — which fields pass, which translate and which drop
