Agent integration guide
One page for an AI agent that builds the gateway into a product: keys, model calls, provider pinning, sessions and caching, errors, limits and cost.
This page is for an AI coding agent that builds RouterPlus into a product's code. It covers the whole integration in order: the key, the SDK, model calls, provider pinning, sessions and prompt caching, errors, limits and cost. Each step links to the reference page that has every detail.
Every docs page is also raw markdown: add .md to its URL, for example https://app.routerplus.com/docs/errors.md. The index of all pages is https://app.routerplus.com/llms.txt. All pages in one file: https://app.routerplus.com/llms-full.txt.
For agents. Do not guess model ids, request fields or headers. Copy them from this page or from the live catalog. If an error message disagrees with your plan, trust the error message and read the page it names. Some steps need a person. They are listed in Steps only a human can do. Stop and ask at those steps.
The short version
- 01Point the OpenAI SDK at
https://api.routerplus.com/v1, or the Anthropic SDK athttps://api.routerplus.com. Read the key fromTM_API_KEY, on the server only. - 02Copy model ids exactly from the catalog. The gateway never substitutes a model: an unknown id is a 404.
- 03Send
max_tokenson every request, from 1 to 32,768. Without it, the gateway uses 4,096. - 04When you need provider features such as prompt caching, call Claude models on
/v1/messagesand other models on/v1/chat/completions. - 05Keep the start of the prompt identical from turn to turn, and only append. This keeps prompt caches warm.
- 06Pin a provider only when you must, with the
providerobject. Check the plan first withPOST /v1/route. - 07Switch on
error_type, never on message text. Retryrate_limitafterretry-after. Stop oninsufficient_quotaand tell a human. - 08There is no idempotency key. A retry is a new request, and it is billed again if it reaches a provider.
- 09Keep at most 8 requests in flight per organization, unless you raised that limit.
- 10Log
x-request-idandusage.costfor every call.
At a glance
| Item | Value |
|---|---|
| OpenAI-compatible base URL | https://api.routerplus.com/v1, for POST /v1/chat/completions |
| Anthropic-compatible base URL | https://api.routerplus.com, for POST /v1/messages (the SDK adds /v1) |
| API key | tm_vk_ followed by 48 hex characters |
| Auth header | Authorization: Bearer <key> or x-api-key: <key>, on every route |
| Model list | GET https://api.routerplus.com/v1/models with your key, or https://app.routerplus.com/api/models.json (public, with prices) |
| Cost of a call | usage.cost, in USD, in every billed response |
| Audit of a call | the x-request-id response header, then GET https://api.routerplus.com/v1/generation?id=<id> |
| Images, video and decisions | POST /v1/images/generations, POST /v1/videos and POST /v1/decisions |
What is not available
Do not build on these. They do not exist on the gateway today:
- The OpenAI Responses API, embeddings, audio, moderation, files and batches. These paths return 404
not_found. Use Chat Completions or Messages. - Server-side conversation memory and idempotency keys. (Session ids exist: see Sessions and caching.)
- More than one answer per request (
nabove 1 is a 400). - Model aliases and model fallback lists. OpenRouter's
modelsarray androutefield are dropped. - Sorting providers by price or speed, and price limits.
provider.sortandprovider.max_priceare a 400. - Provider keys or endpoints in a request.
api_key,base_urlandconnection_idin the body are a 400.
1. Get a key and keep it safe
| Key | Where it comes from | Requests per minute | Use it for |
|---|---|---|---|
| First trial key | Shown after verified browser sign-in through Clerk | 20 | The first test calls |
| Console key | https://app.routerplus.com/console/keys | 300 | Production, staging and CI |
| Identity API key | POST https://app.routerplus.com/api/identity/keys | 300 | One key for each end user or service (step 7) |
To get a key, the user opens Sign up, completes Clerk authentication and email verification, and saves the first key shown after sign-in. Add paid credits in Billing; signup does not grant free credit. Existing customers sign in at https://app.routerplus.com/login using the same verified email and create a key at API keys. Each raw key is shown once.
We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match.
Browser signup is required; the former programmatic signup and local magic-link submission routes have been removed. Once TM_API_KEY is available, the agent can configure clients and make API calls normally.
Rules for the key:
- 01Keep the key on the server. Read it from the
TM_API_KEYenvironment variable or from your secret store. Never commit it, log it or put it in a URL. - 02Do not put the key in a browser or a mobile app. Anyone can copy it from there. The gateway accepts cross-origin requests, so a copied key works from any web page. If a browser-only internal tool must call the gateway, give it its own key with a low monthly spend cap.
- 03Use one key for each service and each environment. Spend, caps and logs are kept per key, and you can disable one key without stopping the others.
- 04Use a console key in production. A trial key allows only 20 requests per minute.
- 05To rotate a key: create the new key, deploy it, then disable the old key in the console. A disabled key stops working within about 5 seconds. The console cannot enable it again. An owner or admin can, with the identity API.
Details: Authentication.
2. Connect the SDK
Change two values in the client you already use: the base URL and the key. Set both in code. Do not depend on OPENAI_BASE_URL or ANTHROPIC_BASE_URL in the environment. When that variable is missing, the SDK sends the request, with its key, to the provider's own API instead of the gateway.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.routerplus.com/v1",
api_key=os.environ["TM_API_KEY"],
# Optional: these name your app in usage analytics. They carry no content.
default_headers={"HTTP-Referer": "https://your-app.example", "X-Title": "Your App"},
)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.routerplus.com/v1",
apiKey: process.env.TM_API_KEY,
defaultHeaders: { "HTTP-Referer": "https://your-app.example", "X-Title": "Your App" },
});With the Anthropic SDK, the base URL has no /v1:
import os
import anthropic
client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ baseURL: "https://api.routerplus.com", apiKey: process.env.TM_API_KEY });Which surface to call
Every chat model answers on both surfaces, because the gateway translates between the two formats. But a request keeps all its provider features only when your surface matches the format of the route that serves it. Then the gateway forwards the request as you sent it. The table gives the native surface of each model family's first route:
| Model ids | Native surface | Features that need it |
|---|---|---|
claude-* | /v1/messages (Anthropic SDK) | Prompt caching with cache_control, thinking, server tools, image input |
gpt-* | /v1/chat/completions (OpenAI SDK) | response_format, reasoning_effort, seed, logprobs, image input |
author/model ids, for example deepseek/deepseek-v4-flash | /v1/chat/completions (OpenAI SDK) | The same OpenAI fields, where the host supports them |
On the other surface, the gateway translates. A field that the translation cannot carry is dropped and named in the x-tm-dropped-params response header. If dropping the field would change the answer, the request is a 400 that names the field instead. Before it refuses, the gateway tries another provider of the same model that speaks your format. The x-tm-provider header names the provider that served you. The catalog gives each model's native format in providers[0].dialect. If your code sends only text and function tools, either surface is fine for every model.
The native surface helps only while the first route serves you. A Claude request on /v1/messages that is pinned to openrouter, or fails over to it, is translated: thinking, server tools and image blocks cannot cross (the request is a 400 if no other route can take it), and other Anthropic-only fields such as top_k, metadata and cache_control are dropped and named.
Details: Wire compatibility.
3. Call models
Rules for every request
- 01Model ids are exact. Copy them from
GET /v1/modelsor/api/models.json. Claude and GPT ids are bare:claude-sonnet-5,gpt-4o-mini. Open models keep their author prefix:moonshotai/kimi-k3. OpenRouter variants such as:freedo not exist here. A dated snapshot of a listed model, such asclaude-haiku-4-5-20251001, also works: it routes and bills as its model. - 02Keep model ids in configuration, in one place. Check them against
GET /v1/modelswhen your service starts. Then a model change is a configuration change. - 03Use the right route for the output. Chat models answer on the two chat surfaces. Image models answer only on
POST /v1/images/generations, video models only onPOST /v1/videos, and decision models only onPOST /v1/decisions. The catalog fieldoutput_modalities(text,image,videoordecisions) tells you which is which. A decision model also takes its own question format: the catalog fielddecisions_wirein/api/models.jsonnames it (systemonefor Jev, Mercury Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider v1 27B and Sage, andgliner), and POST /v1/decisions gives each one. Do not guess the format from the model's name. - 04Send
max_tokensevery time, from 1 to 32,768. If you leave it out, the gateway writesmax_tokens: 4096into the request, and a longer answer stops withfinish_reason: "length". A value above 32,768 is a 400. The gateway also holds balance formax_tokenswhile the call runs, so a realistic value helps on a small balance. - 05Ask for one answer.
nabove 1 is a 400. Make separate calls for several samples. - 06Send the whole conversation on every turn. The gateway keeps no conversation state, and it stores no prompts or answers.
- 07Keep
temperatureat 1 or less in code that can reach Claude models. On a translated request, a higher value is a 400. - 08Refuse drops where they matter. A field outside the forwarded set is dropped and named in
x-tm-dropped-params. Send"provider": {"require_parameters": true}to get a 400 instead of a drop. The 400 comes from the first route that would drop a field: the gateway does not go on to a later route that could carry it. So combine it withonlyororder. For example, theanthropicroute dropsseed, so forseedon a Claude model over/v1/chat/completions, send{"only": ["openrouter"], "require_parameters": true}.
Streaming
Stream answers that a person watches, and long answers. The OpenAI surface sends, in this order: a role chunk, content chunks, a finish chunk, a usage chunk with an empty choices list and usage.cost, then data: [DONE]. The Anthropic surface sends the native event sequence, with usage and cost in the final message_delta.
- If the provider fails after the answer starts, the stream ends with one error event and nothing after it (no
[DONE]). The SDKs raise an error. The partial answer is billed. The gateway never continues an answer on another provider. To try again, send the whole turn again. - During a silence of 15 seconds or more, the gateway sends a keep-alive: an SSE comment on the OpenAI surface, a
pingevent on the Anthropic surface. The SDKs skip them. If you parse SSE yourself, skip them too. - The first response headers must arrive within 20 seconds, or the gateway tries the next provider. A request can run for at most 450 seconds.
- To stop an answer, close the connection. The provider stops. You pay for the usage that the provider reported. If it reported none, you pay for an estimate of the text already streamed, at about four characters per token. If no text streamed yet, the gateway sees no usage, and you pay the full hold for the call: the input estimate plus
max_tokens, at the model's prices.
Python — the OpenAI SDK, keeping the cost and handling a failure in the middle of an answer:
from openai import APIError
parts, usage = [], None
try:
stream = client.chat.completions.create(
model="claude-sonnet-5",
max_tokens=2048,
stream=True,
messages=messages,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
parts.append(chunk.choices[0].delta.content)
if chunk.usage: # the usage chunk; when usage is known, it also comes before an error
usage = chunk.usage
except APIError:
# The answer failed part way. What arrived is billed. Do not join a retry onto
# `parts`: send the whole turn again, or show the error.
raise
print("".join(parts), "cost:", usage.cost if usage else None)Details: Streaming.
Calls that do not stream
Set the client timeout above 450 seconds. The SDK default of 10 minutes is fine. The Anthropic SDKs refuse a call that does not stream when max_tokens is above about 21,000, unless you pass an explicit timeout. Stream those calls instead.
Do not close the connection of a call that does not stream. The gateway stops the provider, but it sees no usage, so you pay the full hold for the call: the input estimate plus max_tokens, at the model's prices. If you may need to cancel, stream the call. An image render is different: it does not stop. It finishes and is billed at the usage that the provider reports.
Tools, reasoning and structured output
- Function tools work on both surfaces for every chat model. The gateway translates tool calls and tool results.
- Anthropic server tools (web search, computer use, bash, text editor) work only for Claude models on
/v1/messages, while theanthropicroute serves them. Theanthropicroute refusesmcp_serversandcontainer. The request then falls over toopenrouter, which runs it without them and names them inx-tm-dropped-params. Send"provider": {"require_parameters": true}to get a 400 instead. thinkingworks only for Claude models on/v1/messages, while theanthropicroute serves them. Reasoning comes back in different fields. On/v1/chat/completions, Claude's thinking arrives asreasoning_content, and open models return OpenRouter's ownreasoningfield (andreasoning_details). On/v1/messages, a streamed answer from a non-Claude model carries its reasoning asthinkingblocks, and a non-streamed answer does not carry it. You can sendthinkingblocks back in the next turn.response_formatneeds an OpenAI-format provider: GPT andauthor/modelids. For JSON from a Claude model, force a tool call withtool_choiceand read the tool input.- Do not end the message list with an assistant message when a Claude model can serve the request from the OpenAI surface. That is a 400, because Claude rejects a prefilled answer.
Images in prompts
Send image parts on the model's native surface: Claude on /v1/messages, GPT and author/model ids on /v1/chat/completions. When the gateway must translate, an image part is a 400 that names the part. Send an image, document, audio or video part only to a model that takes it: architecture.input_modalities in GET /v1/models lists each model's inputs. A part the model does not take is a 400 that names the part, and nothing is billed. Each image or document part counts as 65,536 tokens against your limits and your balance hold while the call runs.
Other routes
- Images:
POST /v1/images/generationstakes the OpenAI Images API body. The image comes back as base64 in the response (response_format: "url"is a 400).nis 1 unless the model allows more: seesupported_parameters.nin/api/models.json(GPT Image models allow up to 4). There is no streaming. A failed render bills nothing. See POST /v1/images/generations. - Video:
POST /v1/videosstarts a job. PollGET /v1/videos/{id}until it ends, then downloadGET /v1/videos/{id}/content. A failed job costs nothing. See POST /v1/videos. - Decisions:
POST /v1/decisionssends a state and typed questions to a decision model:typesafe/jev-1.13(Jev, from TypeSafe),inception/mercury-decide(Mercury Decide, from Inception),bespokelabs/nimble-v3(Bespoke Nimble v3, from Bespoke Labs),cloudflare/clefandcloudflare/clef-flash(Clef and Clef-flash, from Cloudflare),routerplus/decider-2bandrouterplus/kev-4b(Decider 2B and Kev 4B, from RouterPlus),perplexity/pplx-decider-v1-27b(Perplexity Decider v1 27B, from Perplexity),levanto/sage-1.2(Sage, from Levanto) orfastino/gliner-2.5-decide(GLiNER-2.5-Decide, from Fastino). It returns one typed answer for each question, such as a choice, a score or a yes/no probability. The question format follows the model: System One questions (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage) have atype, GLiNER takes aschemaof classifications and extractions in place ofquestions, and a question in another model's format is a 400. GLiNER's numbers are confidences, not calibrated probabilities. Perplexity Decider bills the state once per question, so on it ask only the questions you need. The gateway never fails over from one model to another. Use it for routing, classification and moderation points in your code, in place of a chat model that you parse. See POST /v1/decisions. - Token counting:
POST /v1/messages/count_tokensis free, for models that have an Anthropic-format provider. It still counts against your requests per minute. - Your deployed models: an id such as
tm/<name>-v1from Optimize works on/v1/chat/completionsonly, with keys of the organization that deployed it. Withstream: trueit returns the finished answer as one chunk.
4. Pin providers (only when you must)
How routing works without a pin:
- Each chat model has a fixed, ordered list of providers. Today Claude models go to
anthropicfirst andopenroutersecond. GPT models go toopenaifirst andopenroutersecond. Open models (author/modelids) go toopenrouteronly. - The order does not change from request to request. There is no random spread across providers.
- The gateway moves to the next provider only when the current one fails before it sends any output, or has no capacity left.
- OpenRouter picks the host that runs an open model, for example Fireworks, Together or DeepInfra, for each request.
- The price of a model is the same whichever provider serves it.
Pin when a rule requires one provider, when a feature exists at one provider only, when a session must stay on one host (step 5), or when you compare providers in a test. A pin can cost you fallbacks: only and "allow_fallbacks": false remove the fallback provider, and upstream removes the fallback hosts inside OpenRouter. order keeps every fallback.
The provider object
Send it in the request body on /v1/chat/completions or /v1/messages. The gateway reads it and never forwards it. An unknown key inside it is a 400.
| Field | Value | Effect |
|---|---|---|
order | provider ids | Try these providers first, in this order. The others stay as fallbacks. |
only | provider ids | Use only these providers. |
ignore | provider ids | Never use these providers. |
allow_fallbacks | boolean | false keeps only the first route. If that route is resting after failures, the request fails with 503 and retry-after instead of moving on. |
upstream | host tags, up to 8 | On an OpenRouter route: use only these hosts, in this order, and never another host. If they all fail, the request fails. |
require_parameters | boolean | true: when the first route that would serve you would drop one of your request fields, the whole request is a 400. The gateway does not try a later route. |
connections, connection_order, funding | ids, "byok" or "house" | Choose among your own provider connections and who pays. See Routing policies. |
regions, zdr, data_collection | region list, boolean, "allow" or "deny" | Hard filters. A route passes only with operator evidence, and marketplace routes have none today. So on marketplace traffic, regions, "zdr": true or "data_collection": "deny" removes every route. |
order, only and ignore take route labels and OpenRouter host tags. The marketplace routes are anthropic, openai and openrouter: the providers[].id values in /api/models.json. Your own connections use their profile: openai, anthropic, azure or bedrock.
- A label that matches a route label selects or orders that route, even when that route does not serve the model. Route labels win over host tags with the same spelling:
only: ["anthropic"]means the Anthropic route, so on an open model it leaves no route. - Any other label is a host tag. When the
openrouterroute runs, the gateway sends the tags to OpenRouter as its ownorder,onlyandignore, with yourallow_fallbacks(default true). OpenRouter then prefers, limits or skips those hosts, and it can still fall back to other hosts unlessallow_fallbacksisfalse. - A host tag in
onlykeeps theopenrouterroute open and closes the routes thatonlydoes not name. upstreamis the strict pin: those hosts only, in that order. Host tags fromorder,onlyandignoredo not go with it.POST /v1/routeshows the host preferences in the fieldopenrouter_provider(nullwhen there are none).
Host tags are OpenRouter's provider slugs, for example fireworks, together, deepinfra, amazon-bedrock or google-vertex. The hosts that run a model are its served_by entries in /api/models.json. Use only those hosts: the gateway tells OpenRouter to skip some hosts that OpenRouter lists, and a request that names only those hosts fails. For most hosts, served_by[].id is the same as the tag. Three differ: amazon (tag amazon-bedrock), moonshot (tag moonshotai) and zai (tag z-ai).
| Goal | provider |
|---|---|
| Claude only from Anthropic's own API | {"only": ["anthropic"]} |
| Claude through OpenRouter first, Anthropic as the fallback | {"order": ["openrouter"]} |
| Claude on Amazon Bedrock | {"only": ["openrouter"], "upstream": ["amazon-bedrock"]} |
| An open model on one host for every turn | {"upstream": ["fireworks"]} |
| An open model on one host, with one named backup host | {"upstream": ["fireworks", "together"]} |
| Fail instead of falling back to another provider | add "allow_fallbacks": false |
Check a pin before you ship it. POST /v1/route returns the ordered candidates and each excluded route with its reason. It calls no provider and costs nothing. If your controls exclude every route, a real request returns 404 model_unavailable, the same error as an unknown model. So when a pinned request gets a 404, run /v1/route.
curl -s https://api.routerplus.com/v1/route \
-H "Authorization: Bearer $TM_API_KEY" \
-H "content-type: application/json" \
-d '{"model":"claude-sonnet-5","provider":{"only":["openrouter"],"upstream":["amazon-bedrock"]}}'extra_bodyreply = client.chat.completions.create(
model="moonshotai/kimi-k3",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
extra_body={"provider": {"upstream": ["fireworks"]}},
)const request = {
model: "moonshotai/kimi-k3",
max_tokens: 1024,
messages: [{ role: "user" as const, content: "Hello" }],
provider: { upstream: ["fireworks"] }, // read by the gateway, never forwarded
};
const reply = await client.chat.completions.create(request);Check each answer with its response headers: x-tm-provider names the provider, x-tm-served-by names the host on an OpenRouter route as OpenRouter spells it (for example Amazon Bedrock, not the tag amazon-bedrock), x-tm-upstream-model gives the id the provider received when it differs from yours, and x-tm-attempts counts provider attempts. GET /v1/generation?id=<x-request-id> shows the route of every attempt.
To pin for a whole organization, workspace or key instead of each request, save a routing policy with the policy API. It takes a console session, so a human sets it up. A request's provider object can narrow a saved policy, but it can never add routes. See Routing policies.
Your own provider keys (BYOK)
You can route through your own OpenAI, Anthropic or Azure key, or an AWS Bedrock role. A human adds it once in the console as a connection. Your code still calls the gateway with a marketplace key. The provider bills you directly, and usage.cost is 0. A key bound to a connection can call only that connection's exact model ids, and it does not fall back to marketplace supply unless a routing policy allows it. See Bring your own key.
5. Keep sessions on one route and caches warm
A prompt cache saves money and time only when the next request reaches the same provider, or the same host, with the same prompt start. This is how the gateway treats sessions today:
- The gateway keeps no session state. Your app stores the conversation and sends all of it on every turn.
- Routing is deterministic. For the same model, key and
providerobject, every request tries the same route first. So the turns of a session stay on one provider without a session id. - A turn moves to the next route when the first route fails before it sends output (one failure is enough), when the route's capacity pool is full, or while its provider's 429 pause lasts (the provider's
retry-after, 1 to 60 seconds). After two failures in a row, later turns skip the route for 30 seconds. A turn that moves can miss the cache once. On a follow-up turn, the gateway first waits up to 5 seconds for a full pool: see Waiting for the first route. - Send a session id, and the gateway gives it to each provider in the provider's own field: see Sessions and caching. Claude Code's
x-claude-code-session-idheader counts, with no change on your side. - For an open model, OpenRouter picks the host. With a session id (or the id that the gateway makes when you send none), OpenRouter keeps the conversation on one host from its first request, but it does not promise to. To force one host, pin it with
upstream, as shown below. Remember thatupstreamturns off host fallback: if the pinned hosts fail, the request fails.
Rules for a high cache hit rate:
- 01Keep the start of the prompt byte-identical: the same system prompt, the same tool definitions in the same order, and earlier messages unchanged. Only append.
- 02Put content that changes (the time, ids, documents retrieved for this turn) after the stable part, at the end.
- 03Keep the model and the
providerobject the same for the whole session. - 04For Claude, put
cache_controlon the block that ends the stable part (a system, tool or message block), or send the top-levelcache_controlfield for automatic caching. Both surfaces keep the marks for Anthropic and OpenRouter: see Cache marks. On marketplace routes the cache lasts 5 minutes:ttl: "1h"is removed. - 05OpenAI caches long repeated prompt starts by itself. Many open-model hosts do too. You set nothing for them.
- 06For an open model, pin one host for each session when cache hits matter more than host fallback.
- 07Measure. Cache reads are
usage.prompt_tokens_details.cached_tokenson the OpenAI surface andusage.cache_read_input_tokenson the Anthropic surface. Cache writes arecache_write_tokensandcache_creation_input_tokens. Cache reads are billed at the lowercached_promptprice. Claude cache writes are billed at the highercache_writeprice. Prices are in/api/models.json.
Python — the Anthropic SDK, with a cache mark on a long system prompt that does not change:
reply = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
system=[{
"type": "text",
"text": STABLE_INSTRUCTIONS, # the same bytes on every turn
"cache_control": {"type": "ephemeral"}, # cache everything up to here
}],
messages=history, # append only
)
print(reply.usage) # cache_read_input_tokens, cache_creation_input_tokens and costThe provider caches a prompt start only above a minimum length, so a short system prompt shows no cache reads.
For open models, give each session its own host. The same session always gets the same first host, and different sessions spread across hosts. The second host is used only when the first one fails:
import hashlib
# Hosts that run the model: its served_by entries in https://app.routerplus.com/api/models.json (tags as in step 4)
HOSTS = ["fireworks", "together", "deepinfra"]
def session_route(session_id: str) -> dict:
i = int(hashlib.sha256(session_id.encode()).hexdigest(), 16) % len(HOSTS)
return {"upstream": [HOSTS[i], HOSTS[(i + 1) % len(HOSTS)]]}
reply = client.chat.completions.create(
model="moonshotai/kimi-k3",
max_tokens=2048,
messages=history,
extra_body={"provider": session_route(session_id)},
)
print(reply.usage.prompt_tokens_details.cached_tokens)Sessions and caching
Session ids
On POST /v1/chat/completions and POST /v1/messages, the session id is the first valid value among these inputs, in this order. Token counting, images and videos do not use it.
| Order | Input | Kind |
|---|---|---|
| 1 | session_id | Top-level body field |
| 2 | x-session-id | Header |
| 3 | x-session-affinity | Header |
| 4 | x-claude-code-session-id | Header that Claude Code sends by itself |
- A valid value has 1 to 256 printable ASCII characters.
prompt_cache_key, a top-level body field, follows the same rule. - The gateway ignores an invalid value and uses the next input. An invalid value alone is never an error. The ignored input is named in
x-tm-dropped-paramsassession_id (invalid),prompt_cache_key (invalid),header:x-session-id (invalid),header:x-session-affinity (invalid)orheader:x-claude-code-session-id (invalid). With"require_parameters": true, that makes the request a 400, as any dropped field does. - Valid session inputs are never named in
x-tm-dropped-params, and they never triggerrequire_parameters. prompt_cache_keystays OpenAI's own field. It wins over the session id on OpenAI routes, and it is the fallback on OpenRouter.- When two valid inputs disagree, the order decides. There is no error.
- Codex's
session-idheader is not read. It can matter only after the gateway addsPOST /v1/responses, because current Codex calls only the Responses API.
What each route receives:
| Route | Field | Value |
|---|---|---|
openai (marketplace), and openai and azure (your connections) | prompt_cache_key | Your prompt_cache_key, or else the session id. Nothing if you sent neither |
openrouter (marketplace) | session_id | The session id, or else your prompt_cache_key, or else an id that the gateway makes. If you sent prompt_cache_key, it goes too |
anthropic and bedrock | Nothing | Anthropic has no session input. Its cache matches the prompt start |
- Values go to the provider unchanged. A value longer than the provider accepts is replaced by
tm-followed by 43 base64url characters: a hash of the value. - For OpenRouter only, when you send neither a session id nor
prompt_cache_key, the gateway makes an id from your organization id, the model and the opening messages, up to the first user message. OpenRouter then keeps the conversation on one host from its first request. - The gateway keeps no session table. The id goes to the provider, and the provider routes on it.
prompt_cache_retentiongoes unchanged toopenaiandazureroutes, both"in_memory"and"24h". On other routes it is dropped and named.
End-user ids on marketplace routes
On marketplace routes, the provider gets a hash in place of your end user's id, and the raw values do not go. The id is the first string among the body fields safety_identifier, user and metadata.user_id. The hash is tm- followed by 43 base64url characters, made from your organization id and that id. Without an id, it is made from your organization id alone. Each marketplace route gets exactly one field:
| Marketplace route | Field that carries the hash | Fields removed |
|---|---|---|
openai, azure | safety_identifier | user |
openrouter | user | safety_identifier |
anthropic | metadata.user_id | user, safety_identifier |
| every other route (OpenAI-compatible), dedicated capacity included | user | safety_identifier, metadata.user_id |
- A provider that blocks an id then blocks one user of one customer, not the whole marketplace account.
- On your own connections,
userandmetadatapass unchanged.safety_identifierpasses unchanged toopenaiandazureconnections, and it is dropped and named onanthropicandbedrockconnections. - Identity fields that a route uses are never named in
x-tm-dropped-params.
Cache marks
The gateway never adds a cache mark that you did not send. Block and part marks (cache_control) are handled this way:
| Request | Route | What happens to the marks |
|---|---|---|
/v1/chat/completions | Anthropic direct or Bedrock | Each text part keeps its mark. A marked system part turns the system prompt into blocks. A tool message's last mark goes on its tool_result block. Other part keys are not forwarded |
/v1/chat/completions | OpenRouter | Parts pass unchanged |
/v1/chat/completions | OpenAI direct or Azure | Part marks pass unchanged. OpenAI and Azure ignore them |
/v1/messages | Anthropic direct or Bedrock | Block marks pass unchanged |
/v1/messages | OpenRouter | Marks survive in system, user, assistant and tool content. Marks on tool definitions and on tool_use blocks are removed and named as tools[i].cache_control or messages[i].content[j].cache_control |
/v1/messages | OpenAI direct or Azure | Marks are removed and named |
A top-level cache_control field (automatic caching), on either surface:
- Anthropic direct and OpenRouter: sent unchanged.
- Bedrock: turned into a mark on the last block that can carry one, counted back from the end of the messages. If no block can carry it, it is dropped and named as
cache_control. - OpenAI direct and Azure: dropped and named as
cache_control. - Any route: a field that Anthropic would refuse is dropped and named as
cache_control. Anthropic refuses a field other than{"type": "ephemeral"}with an optionalttlof"5m"or"1h", a fifth mark (four marks exist already), a 1-hour field after a 5-minute mark, and a field whose TTL differs from the mark on the last block.
On marketplace routes, ttl: "1h" is removed from every mark, so the cache lasts 5 minutes. The drop is named as <path>.cache_control.ttl, for example system[0].cache_control.ttl, messages[2].content[0].cache_control.ttl, or cache_control.ttl for the top-level field. ttl: "5m" passes. Your own connections keep ttl: "1h".
Waiting for the first route
On a follow-up turn (an assistant or tool message after a user message), if the first route's capacity pool is full, the gateway waits for that pool for up to 5 seconds (an operator setting) before it moves the request to the next route. It never waits past the request's deadline. The x-tm-affinity-wait-ms response header gives the wait in milliseconds, only when the wait was more than 0.
6. Handle errors and retries
Every failure carries one stable class, error_type. Read it from error.metadata.error_type on the OpenAI surface, or error.error_type on the Anthropic surface. The x-tm-error-code header carries it too, but browsers cannot read that header. Never parse the message text. Log the x-request-id header with every failure.
error_type | HTTP | Retry | What to do |
|---|---|---|---|
rate_limit | 429 | Yes, once | Wait retry-after seconds. If x-tm-limit-kind is concurrency, send fewer requests at once. |
email_not_verified | 403 | No | Stop and tell a human to open the verify link in their email; the key starts working within seconds of the click. |
insufficient_quota | 429 | No | Stop and tell a human: add credits or raise a spend cap. On a cap, x-tm-cap-reset gives the reset time. |
upstream_error | 5xx | Yes, once, after about 2 seconds | The gateway already tried every other provider it could. If the retry fails too, report it with the x-request-id. |
upstream_error | 4xx | No | The provider refused the request itself. Read the message. |
gateway_error | 503 | Yes | Wait retry-after. The request did not reach a provider. |
gateway_error | 504 | Yes, once | The deadline passed while the gateway tried providers. An earlier attempt may be billed: check GET /v1/generation?id=<x-request-id> first. |
gateway_error | 400 | No | Send one output limit, from 1 to 32,768. |
gateway_error | 500 | No | Report it with the x-request-id. |
model_unavailable | 404 | No | Fix the model id, or fix your provider object (check it with POST /v1/route). |
model_unavailable | 502 | Yes, once | The connection to every provider failed. |
invalid_request | 400 | No | Fix the field that the message names. |
context_overflow | 400 | No | Shorten the input, or use a model with a larger context. |
content_policy | 400 | No | Change the prompt. It is never sent to another provider. |
auth | 401 | No | Check the header, the key, and whether the key was disabled. A 401, 402 or 403 relayed from every provider is the marketplace's own account problem: report it with the x-request-id. |
request_too_large | 413 | No | Keep the body under 10 MB (1 MB for a video request). |
not_found | 404 | No | Use a route that exists. The list is in Errors. |
Rules:
- 01Let the SDK retry. The official OpenAI and Anthropic SDKs retry 429 and 5xx responses twice by default, and they wait for
retry-after. Keep that default. Do not add a second retry loop around it. The SDKs also retry aninsufficient_quotaresponse. That costs nothing, but your code must then stop. - 02There is no idempotency key. A retried request is a new request. If it reaches a provider, it is billed again. A refusal by the gateway's own limits (
x-tm-error-origin: gateway_admission) cost nothing. - 03Never retry
insufficient_quota,invalid_request,context_overflow,content_policyorauthin a loop. The request, or the account, must change first. - 04An error in the middle of a stream is final for that answer. Send the whole turn again. Never join two partial answers.
Python — the OpenAI SDK, reading the class:
from openai import APIStatusError
try:
reply = client.chat.completions.create(model=MODEL, max_tokens=1024, messages=messages)
except APIStatusError as e:
error = e.response.json().get("error") or {}
kind = (error.get("metadata") or {}).get("error_type") or e.response.headers.get("x-tm-error-code")
request_id = e.response.headers.get("x-request-id")
if kind == "insufficient_quota":
notify_operator(request_id, error.get("message")) # money or a cap: a retry cannot help
raiseDetails: Errors.
7. Limits, spend caps and your users
Default limits:
| Scope | Requests per minute | Tokens per minute | In flight at once |
|---|---|---|---|
| Trial key | 20 | 1,000,000 | 8 |
| Console or identity API key | 300 | 1,000,000 | 8 |
| Organization, all keys together | 300 | 1,000,000 | 8 |
Rules:
- 01Limit how many requests your code sends at once. By default the organization allows 8 in flight. A ninth gets 429
rate_limitwithx-tm-limit-kind: concurrency. Use a semaphore or a fixed pool of workers. To run more at once, a human first raises the organization and key limits in /console/limits. A key's requests per minute cannot go above its ceiling (300 for a console key). Each marketplace provider pool also limits one organization to half of the pool's capacity. That refusal carriesx-tm-limit-scope: pool. - 02Read
x-tm-remaining-rpmandx-tm-remaining-tpmon each response, and slow down before they reach zero.GET https://api.routerplus.com/v1/limits?model=<id>returns the limits that apply to your key. - 03While a call runs, token limits count an estimate: one token for each four bytes of the request, rounded up, 65,536 for each image or document part, plus
max_tokens. An image sent inline as base64 counts its 65,536 only, not its bytes. The real count replaces the estimate when the call ends. Small requests and a realisticmax_tokenslet more calls run at once. - 04Spend caps are monthly, for each key and for the organization, on the UTC calendar month. A human sets them in the console. A cap refuses with 429
insufficient_quotaand thex-tm-cap-resetheader. Give each feature or customer its own key and cap, so a bug or a loop cannot spend everything.
Your users:
- By default, one backend key serves all your users. The gateway does not know who your users are, so enforce per-user quotas in your app.
safety_identifier,userandmetadata.user_ididentify your end user to the provider. On marketplace routes the gateway sends only a hash of that id with your organization; on your own connections they go unchanged (see End-user ids on marketplace routes). They do not identify, limit or bill anyone on the gateway. Do not put emails or names in them.- For limits that the gateway enforces per user, use the identity API. Create an
end_userprincipal for each user, issue a key for it, and set limits on the principal. Your backend keeps those keys. Never give a key to the user. The identity API takes a console session, not an API key, so a human with an owner or admin role sets it up. See Workspaces and principals. - To name your app in analytics, send the
HTTP-RefererandX-Titleheaders. They carry no content.
Details: Rate limits and spend caps and Limits and capacity.
8. Track cost
- Every billed response has
usage.costin USD. It is the full charge for the request, including provider attempts that failed before the answer. Store it with thex-request-id, the model and your own ids, such as the user and the feature. - In a stream,
costis in the usage chunk on the OpenAI surface, and in the lastmessage_deltaon the Anthropic surface. - An Anthropic-surface stream that fails in the middle of the answer has no usage in it. Read its cost from
GET /v1/generation?id=<x-request-id>. GET /v1/usagereturns the balance and the last 20 attempts. For a key bound to its own workspace or principal, the balance fields arenull. The console shows usage and logs for the whole organization.- BYOK calls show a cost of 0, because your provider bills you directly.
- To check a charge, multiply the token counts by the prices in
/api/models.json. The gateway rounds the final charge down.
Details: Pricing and billing and GET /v1/usage and /v1/generation.
9. Verify the integration
One cheap call shows the answer, the cost and the headers that the gateway adds:
curl -sS -i https://api.routerplus.com/v1/chat/completions \
-H "Authorization: Bearer $TM_API_KEY" \
-H "content-type: application/json" \
-d '{"model":"gpt-4o-mini","max_tokens":16,"messages":[{"role":"user","content":"Reply with OK"}]}'Then confirm that the ledger saw it:
curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usageThe integration is done when all of these are true:
- 01Every client uses the gateway base URL and reads the key from
TM_API_KEYon the server. - 02The key is not in the repository, in a client bundle or in a log.
- 03Every model id in the configuration appears in
GET /v1/models. - 04Every request sends
max_tokens. - 05One streamed and one non-streamed test call return
usage.cost. - 06The
x-tm-dropped-paramsheader is absent, or it names only fields that you accept losing. - 07
x-tm-provider, andx-tm-served-bywhen you pin a host, show the route that you expect. (x-tm-served-byuses OpenRouter's display name, not the tag.) - 08The error handling switches on
error_type. It stops oninsufficient_quota, and it adds no retry loop on top of the SDK's own retries. - 09The number of requests in flight stays within the organization's limit.
- 10The logs keep
x-request-idandusage.costfor every call. - 11The test calls appear in
GET /v1/usage.
Leave instructions for the next agent
Add this block to the AGENTS.md or CLAUDE.md file of the repository that you integrated. Later coding sessions then keep the integration correct:
## LLM calls: RouterPlus gateway
- Every LLM call goes through the RouterPlus gateway. Full guide: https://app.routerplus.com/docs/agent-integration.md
- OpenAI SDK base URL: https://api.routerplus.com/v1. Anthropic SDK base URL: https://api.routerplus.com. Key: the TM_API_KEY environment variable, on the server only.
- Model ids live in configuration. Each one must exist in https://app.routerplus.com/api/models.json. Never invent or alias an id.
- Send max_tokens on every request (1 to 32768). Never send n above 1.
- Claude models: /v1/messages, with cache_control on the stable start of the prompt. GPT and author/model ids: /v1/chat/completions.
- Keep the start of the prompt byte-identical across turns. Only append.
- Send one id per conversation as the x-session-id header (or session_id in the body). Providers then keep its cache warm.
- Provider pins go in the request body "provider" object (only, ignore, order, allow_fallbacks, upstream, require_parameters). Check a pin with POST https://api.routerplus.com/v1/route.
- Errors: switch on error_type. rate_limit: wait retry-after, then retry once. insufficient_quota: stop and tell a human. Never retry invalid_request, context_overflow or content_policy.
- No idempotency key: a retry is a new, billed request. Add no retry loop on top of the SDK's own retries.
- At most 8 requests in flight per organization, unless the limit was raised.
- Log x-request-id and usage.cost for every call.
- Not available: the Responses API, embeddings, audio, files and batches.Steps only a human can do
- Complete browser signup through Clerk, including email verification, then buy credits in Billing. Signup and card verification alone do not add credit.
- Fix billing when
insufficient_quotaappears: add credits, or raise a spend cap in the console. - Sign in to the console to create keys, set spend caps and change limits.
- Add your own provider keys (BYOK), save routing policies, and set up workspaces and end-user principals.
- Approve large test runs. Every call is billed.
Where next
Markdown source for agents: /docs/agent-integration.md · index at /llms.txt