Console
Get started/Agent integration guide

Agent integration guide

One page for an AI agent that builds the gateway into a product: keys, model calls, provider pinning, sessions and caching, errors, limits and cost.

/llms.txt

This page is for an AI coding agent that builds RouterPlus into a product's code. It covers the whole integration in order: the key, the SDK, model calls, provider pinning, sessions and prompt caching, errors, limits and cost. Each step links to the reference page that has every detail.

Every docs page is also raw markdown: add .md to its URL, for example https://app.routerplus.com/docs/errors.md. The index of all pages is https://app.routerplus.com/llms.txt. All pages in one file: https://app.routerplus.com/llms-full.txt.

Note

For agents. Do not guess model ids, request fields or headers. Copy them from this page or from the live catalog. If an error message disagrees with your plan, trust the error message and read the page it names. Some steps need a person. They are listed in Steps only a human can do. Stop and ask at those steps.

The short version

  1. 01
    Point the OpenAI SDK at https://api.routerplus.com/v1, or the Anthropic SDK at https://api.routerplus.com. Read the key from TM_API_KEY, on the server only.
  2. 02
    Copy model ids exactly from the catalog. The gateway never substitutes a model: an unknown id is a 404.
  3. 03
    Send max_tokens on every request, from 1 to 32,768. Without it, the gateway uses 4,096.
  4. 04
    When you need provider features such as prompt caching, call Claude models on /v1/messages and other models on /v1/chat/completions.
  5. 05
    Keep the start of the prompt identical from turn to turn, and only append. This keeps prompt caches warm.
  6. 06
    Pin a provider only when you must, with the provider object. Check the plan first with POST /v1/route.
  7. 07
    Switch on error_type, never on message text. Retry rate_limit after retry-after. Stop on insufficient_quota and tell a human.
  8. 08
    There is no idempotency key. A retry is a new request, and it is billed again if it reaches a provider.
  9. 09
    Keep at most 8 requests in flight per organization, unless you raised that limit.
  10. 10
    Log x-request-id and usage.cost for every call.

At a glance

ItemValue
OpenAI-compatible base URLhttps://api.routerplus.com/v1, for POST /v1/chat/completions
Anthropic-compatible base URLhttps://api.routerplus.com, for POST /v1/messages (the SDK adds /v1)
API keytm_vk_ followed by 48 hex characters
Auth headerAuthorization: Bearer <key> or x-api-key: <key>, on every route
Model listGET https://api.routerplus.com/v1/models with your key, or https://app.routerplus.com/api/models.json (public, with prices)
Cost of a callusage.cost, in USD, in every billed response
Audit of a callthe x-request-id response header, then GET https://api.routerplus.com/v1/generation?id=<id>
Images, video and decisionsPOST /v1/images/generations, POST /v1/videos and POST /v1/decisions

What is not available

Do not build on these. They do not exist on the gateway today:

  • The OpenAI Responses API, embeddings, audio, moderation, files and batches. These paths return 404 not_found. Use Chat Completions or Messages.
  • Server-side conversation memory and idempotency keys. (Session ids exist: see Sessions and caching.)
  • More than one answer per request (n above 1 is a 400).
  • Model aliases and model fallback lists. OpenRouter's models array and route field are dropped.
  • Sorting providers by price or speed, and price limits. provider.sort and provider.max_price are a 400.
  • Provider keys or endpoints in a request. api_key, base_url and connection_id in the body are a 400.

1. Get a key and keep it safe

KeyWhere it comes fromRequests per minuteUse it for
First trial keyShown after verified browser sign-in through Clerk20The first test calls
Console keyhttps://app.routerplus.com/console/keys300Production, staging and CI
Identity API keyPOST https://app.routerplus.com/api/identity/keys300One key for each end user or service (step 7)

To get a key, the user opens Sign up, completes Clerk authentication and email verification, and saves the first key shown after sign-in. Add paid credits in Billing; signup does not grant free credit. Existing customers sign in at https://app.routerplus.com/login using the same verified email and create a key at API keys. Each raw key is shown once.

We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match.

Browser signup is required; the former programmatic signup and local magic-link submission routes have been removed. Once TM_API_KEY is available, the agent can configure clients and make API calls normally.

Rules for the key:

  1. 01
    Keep the key on the server. Read it from the TM_API_KEY environment variable or from your secret store. Never commit it, log it or put it in a URL.
  2. 02
    Do not put the key in a browser or a mobile app. Anyone can copy it from there. The gateway accepts cross-origin requests, so a copied key works from any web page. If a browser-only internal tool must call the gateway, give it its own key with a low monthly spend cap.
  3. 03
    Use one key for each service and each environment. Spend, caps and logs are kept per key, and you can disable one key without stopping the others.
  4. 04
    Use a console key in production. A trial key allows only 20 requests per minute.
  5. 05
    To rotate a key: create the new key, deploy it, then disable the old key in the console. A disabled key stops working within about 5 seconds. The console cannot enable it again. An owner or admin can, with the identity API.

Details: Authentication.

2. Connect the SDK

Change two values in the client you already use: the base URL and the key. Set both in code. Do not depend on OPENAI_BASE_URL or ANTHROPIC_BASE_URL in the environment. When that variable is missing, the SDK sends the request, with its key, to the provider's own API instead of the gateway.

python
The OpenAI SDK
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.routerplus.com/v1",
    api_key=os.environ["TM_API_KEY"],
    # Optional: these name your app in usage analytics. They carry no content.
    default_headers={"HTTP-Referer": "https://your-app.example", "X-Title": "Your App"},
)
typescript
The OpenAI SDK
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.routerplus.com/v1",
  apiKey: process.env.TM_API_KEY,
  defaultHeaders: { "HTTP-Referer": "https://your-app.example", "X-Title": "Your App" },
});

With the Anthropic SDK, the base URL has no /v1:

python
import os
import anthropic

client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])
typescript
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({ baseURL: "https://api.routerplus.com", apiKey: process.env.TM_API_KEY });

Which surface to call

Every chat model answers on both surfaces, because the gateway translates between the two formats. But a request keeps all its provider features only when your surface matches the format of the route that serves it. Then the gateway forwards the request as you sent it. The table gives the native surface of each model family's first route:

Model idsNative surfaceFeatures that need it
claude-*/v1/messages (Anthropic SDK)Prompt caching with cache_control, thinking, server tools, image input
gpt-*/v1/chat/completions (OpenAI SDK)response_format, reasoning_effort, seed, logprobs, image input
author/model ids, for example deepseek/deepseek-v4-flash/v1/chat/completions (OpenAI SDK)The same OpenAI fields, where the host supports them

On the other surface, the gateway translates. A field that the translation cannot carry is dropped and named in the x-tm-dropped-params response header. If dropping the field would change the answer, the request is a 400 that names the field instead. Before it refuses, the gateway tries another provider of the same model that speaks your format. The x-tm-provider header names the provider that served you. The catalog gives each model's native format in providers[0].dialect. If your code sends only text and function tools, either surface is fine for every model.

The native surface helps only while the first route serves you. A Claude request on /v1/messages that is pinned to openrouter, or fails over to it, is translated: thinking, server tools and image blocks cannot cross (the request is a 400 if no other route can take it), and other Anthropic-only fields such as top_k, metadata and cache_control are dropped and named.

Details: Wire compatibility.

3. Call models

Rules for every request

  1. 01
    Model ids are exact. Copy them from GET /v1/models or /api/models.json. Claude and GPT ids are bare: claude-sonnet-5, gpt-4o-mini. Open models keep their author prefix: moonshotai/kimi-k3. OpenRouter variants such as :free do not exist here. A dated snapshot of a listed model, such as claude-haiku-4-5-20251001, also works: it routes and bills as its model.
  2. 02
    Keep model ids in configuration, in one place. Check them against GET /v1/models when your service starts. Then a model change is a configuration change.
  3. 03
    Use the right route for the output. Chat models answer on the two chat surfaces. Image models answer only on POST /v1/images/generations, video models only on POST /v1/videos, and decision models only on POST /v1/decisions. The catalog field output_modalities (text, image, video or decisions) tells you which is which. A decision model also takes its own question format: the catalog field decisions_wire in /api/models.json names it (systemone for Jev, Mercury Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider v1 27B and Sage, and gliner), and POST /v1/decisions gives each one. Do not guess the format from the model's name.
  4. 04
    Send max_tokens every time, from 1 to 32,768. If you leave it out, the gateway writes max_tokens: 4096 into the request, and a longer answer stops with finish_reason: "length". A value above 32,768 is a 400. The gateway also holds balance for max_tokens while the call runs, so a realistic value helps on a small balance.
  5. 05
    Ask for one answer. n above 1 is a 400. Make separate calls for several samples.
  6. 06
    Send the whole conversation on every turn. The gateway keeps no conversation state, and it stores no prompts or answers.
  7. 07
    Keep temperature at 1 or less in code that can reach Claude models. On a translated request, a higher value is a 400.
  8. 08
    Refuse drops where they matter. A field outside the forwarded set is dropped and named in x-tm-dropped-params. Send "provider": {"require_parameters": true} to get a 400 instead of a drop. The 400 comes from the first route that would drop a field: the gateway does not go on to a later route that could carry it. So combine it with only or order. For example, the anthropic route drops seed, so for seed on a Claude model over /v1/chat/completions, send {"only": ["openrouter"], "require_parameters": true}.

Streaming

Stream answers that a person watches, and long answers. The OpenAI surface sends, in this order: a role chunk, content chunks, a finish chunk, a usage chunk with an empty choices list and usage.cost, then data: [DONE]. The Anthropic surface sends the native event sequence, with usage and cost in the final message_delta.

  • If the provider fails after the answer starts, the stream ends with one error event and nothing after it (no [DONE]). The SDKs raise an error. The partial answer is billed. The gateway never continues an answer on another provider. To try again, send the whole turn again.
  • During a silence of 15 seconds or more, the gateway sends a keep-alive: an SSE comment on the OpenAI surface, a ping event on the Anthropic surface. The SDKs skip them. If you parse SSE yourself, skip them too.
  • The first response headers must arrive within 20 seconds, or the gateway tries the next provider. A request can run for at most 450 seconds.
  • To stop an answer, close the connection. The provider stops. You pay for the usage that the provider reported. If it reported none, you pay for an estimate of the text already streamed, at about four characters per token. If no text streamed yet, the gateway sees no usage, and you pay the full hold for the call: the input estimate plus max_tokens, at the model's prices.

Python — the OpenAI SDK, keeping the cost and handling a failure in the middle of an answer:

python
from openai import APIError

parts, usage = [], None
try:
    stream = client.chat.completions.create(
        model="claude-sonnet-5",
        max_tokens=2048,
        stream=True,
        messages=messages,
    )
    for chunk in stream:
        if chunk.choices and chunk.choices[0].delta.content:
            parts.append(chunk.choices[0].delta.content)
        if chunk.usage:  # the usage chunk; when usage is known, it also comes before an error
            usage = chunk.usage
except APIError:
    # The answer failed part way. What arrived is billed. Do not join a retry onto
    # `parts`: send the whole turn again, or show the error.
    raise
print("".join(parts), "cost:", usage.cost if usage else None)

Details: Streaming.

Calls that do not stream

Set the client timeout above 450 seconds. The SDK default of 10 minutes is fine. The Anthropic SDKs refuse a call that does not stream when max_tokens is above about 21,000, unless you pass an explicit timeout. Stream those calls instead.

Do not close the connection of a call that does not stream. The gateway stops the provider, but it sees no usage, so you pay the full hold for the call: the input estimate plus max_tokens, at the model's prices. If you may need to cancel, stream the call. An image render is different: it does not stop. It finishes and is billed at the usage that the provider reports.

Tools, reasoning and structured output

  • Function tools work on both surfaces for every chat model. The gateway translates tool calls and tool results.
  • Anthropic server tools (web search, computer use, bash, text editor) work only for Claude models on /v1/messages, while the anthropic route serves them. The anthropic route refuses mcp_servers and container. The request then falls over to openrouter, which runs it without them and names them in x-tm-dropped-params. Send "provider": {"require_parameters": true} to get a 400 instead.
  • thinking works only for Claude models on /v1/messages, while the anthropic route serves them. Reasoning comes back in different fields. On /v1/chat/completions, Claude's thinking arrives as reasoning_content, and open models return OpenRouter's own reasoning field (and reasoning_details). On /v1/messages, a streamed answer from a non-Claude model carries its reasoning as thinking blocks, and a non-streamed answer does not carry it. You can send thinking blocks back in the next turn.
  • response_format needs an OpenAI-format provider: GPT and author/model ids. For JSON from a Claude model, force a tool call with tool_choice and read the tool input.
  • Do not end the message list with an assistant message when a Claude model can serve the request from the OpenAI surface. That is a 400, because Claude rejects a prefilled answer.

Images in prompts

Send image parts on the model's native surface: Claude on /v1/messages, GPT and author/model ids on /v1/chat/completions. When the gateway must translate, an image part is a 400 that names the part. Send an image, document, audio or video part only to a model that takes it: architecture.input_modalities in GET /v1/models lists each model's inputs. A part the model does not take is a 400 that names the part, and nothing is billed. Each image or document part counts as 65,536 tokens against your limits and your balance hold while the call runs.

Other routes

  • Images: POST /v1/images/generations takes the OpenAI Images API body. The image comes back as base64 in the response (response_format: "url" is a 400). n is 1 unless the model allows more: see supported_parameters.n in /api/models.json (GPT Image models allow up to 4). There is no streaming. A failed render bills nothing. See POST /v1/images/generations.
  • Video: POST /v1/videos starts a job. Poll GET /v1/videos/{id} until it ends, then download GET /v1/videos/{id}/content. A failed job costs nothing. See POST /v1/videos.
  • Decisions: POST /v1/decisions sends a state and typed questions to a decision model: typesafe/jev-1.13 (Jev, from TypeSafe), inception/mercury-decide (Mercury Decide, from Inception), bespokelabs/nimble-v3 (Bespoke Nimble v3, from Bespoke Labs), cloudflare/clef and cloudflare/clef-flash (Clef and Clef-flash, from Cloudflare), routerplus/decider-2b and routerplus/kev-4b (Decider 2B and Kev 4B, from RouterPlus), perplexity/pplx-decider-v1-27b (Perplexity Decider v1 27B, from Perplexity), levanto/sage-1.2 (Sage, from Levanto) or fastino/gliner-2.5-decide (GLiNER-2.5-Decide, from Fastino). It returns one typed answer for each question, such as a choice, a score or a yes/no probability. The question format follows the model: System One questions (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage) have a type, GLiNER takes a schema of classifications and extractions in place of questions, and a question in another model's format is a 400. GLiNER's numbers are confidences, not calibrated probabilities. Perplexity Decider bills the state once per question, so on it ask only the questions you need. The gateway never fails over from one model to another. Use it for routing, classification and moderation points in your code, in place of a chat model that you parse. See POST /v1/decisions.
  • Token counting: POST /v1/messages/count_tokens is free, for models that have an Anthropic-format provider. It still counts against your requests per minute.
  • Your deployed models: an id such as tm/<name>-v1 from Optimize works on /v1/chat/completions only, with keys of the organization that deployed it. With stream: true it returns the finished answer as one chunk.

4. Pin providers (only when you must)

How routing works without a pin:

  • Each chat model has a fixed, ordered list of providers. Today Claude models go to anthropic first and openrouter second. GPT models go to openai first and openrouter second. Open models (author/model ids) go to openrouter only.
  • The order does not change from request to request. There is no random spread across providers.
  • The gateway moves to the next provider only when the current one fails before it sends any output, or has no capacity left.
  • OpenRouter picks the host that runs an open model, for example Fireworks, Together or DeepInfra, for each request.
  • The price of a model is the same whichever provider serves it.

Pin when a rule requires one provider, when a feature exists at one provider only, when a session must stay on one host (step 5), or when you compare providers in a test. A pin can cost you fallbacks: only and "allow_fallbacks": false remove the fallback provider, and upstream removes the fallback hosts inside OpenRouter. order keeps every fallback.

The provider object

Send it in the request body on /v1/chat/completions or /v1/messages. The gateway reads it and never forwards it. An unknown key inside it is a 400.

FieldValueEffect
orderprovider idsTry these providers first, in this order. The others stay as fallbacks.
onlyprovider idsUse only these providers.
ignoreprovider idsNever use these providers.
allow_fallbacksbooleanfalse keeps only the first route. If that route is resting after failures, the request fails with 503 and retry-after instead of moving on.
upstreamhost tags, up to 8On an OpenRouter route: use only these hosts, in this order, and never another host. If they all fail, the request fails.
require_parametersbooleantrue: when the first route that would serve you would drop one of your request fields, the whole request is a 400. The gateway does not try a later route.
connections, connection_order, fundingids, "byok" or "house"Choose among your own provider connections and who pays. See Routing policies.
regions, zdr, data_collectionregion list, boolean, "allow" or "deny"Hard filters. A route passes only with operator evidence, and marketplace routes have none today. So on marketplace traffic, regions, "zdr": true or "data_collection": "deny" removes every route.

order, only and ignore take route labels and OpenRouter host tags. The marketplace routes are anthropic, openai and openrouter: the providers[].id values in /api/models.json. Your own connections use their profile: openai, anthropic, azure or bedrock.

  • A label that matches a route label selects or orders that route, even when that route does not serve the model. Route labels win over host tags with the same spelling: only: ["anthropic"] means the Anthropic route, so on an open model it leaves no route.
  • Any other label is a host tag. When the openrouter route runs, the gateway sends the tags to OpenRouter as its own order, only and ignore, with your allow_fallbacks (default true). OpenRouter then prefers, limits or skips those hosts, and it can still fall back to other hosts unless allow_fallbacks is false.
  • A host tag in only keeps the openrouter route open and closes the routes that only does not name.
  • upstream is the strict pin: those hosts only, in that order. Host tags from order, only and ignore do not go with it.
  • POST /v1/route shows the host preferences in the field openrouter_provider (null when there are none).

Host tags are OpenRouter's provider slugs, for example fireworks, together, deepinfra, amazon-bedrock or google-vertex. The hosts that run a model are its served_by entries in /api/models.json. Use only those hosts: the gateway tells OpenRouter to skip some hosts that OpenRouter lists, and a request that names only those hosts fails. For most hosts, served_by[].id is the same as the tag. Three differ: amazon (tag amazon-bedrock), moonshot (tag moonshotai) and zai (tag z-ai).

Goalprovider
Claude only from Anthropic's own API{"only": ["anthropic"]}
Claude through OpenRouter first, Anthropic as the fallback{"order": ["openrouter"]}
Claude on Amazon Bedrock{"only": ["openrouter"], "upstream": ["amazon-bedrock"]}
An open model on one host for every turn{"upstream": ["fireworks"]}
An open model on one host, with one named backup host{"upstream": ["fireworks", "together"]}
Fail instead of falling back to another provideradd "allow_fallbacks": false

Check a pin before you ship it. POST /v1/route returns the ordered candidates and each excluded route with its reason. It calls no provider and costs nothing. If your controls exclude every route, a real request returns 404 model_unavailable, the same error as an unknown model. So when a pinned request gets a 404, run /v1/route.

curl
curl -s https://api.routerplus.com/v1/route \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5","provider":{"only":["openrouter"],"upstream":["amazon-bedrock"]}}'
python
The OpenAI SDK sends fields it does not know through extra_body
reply = client.chat.completions.create(
    model="moonshotai/kimi-k3",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}],
    extra_body={"provider": {"upstream": ["fireworks"]}},
)
typescript
Pass a variable, so the field the SDK types do not know is kept
const request = {
  model: "moonshotai/kimi-k3",
  max_tokens: 1024,
  messages: [{ role: "user" as const, content: "Hello" }],
  provider: { upstream: ["fireworks"] }, // read by the gateway, never forwarded
};
const reply = await client.chat.completions.create(request);

Check each answer with its response headers: x-tm-provider names the provider, x-tm-served-by names the host on an OpenRouter route as OpenRouter spells it (for example Amazon Bedrock, not the tag amazon-bedrock), x-tm-upstream-model gives the id the provider received when it differs from yours, and x-tm-attempts counts provider attempts. GET /v1/generation?id=<x-request-id> shows the route of every attempt.

To pin for a whole organization, workspace or key instead of each request, save a routing policy with the policy API. It takes a console session, so a human sets it up. A request's provider object can narrow a saved policy, but it can never add routes. See Routing policies.

Your own provider keys (BYOK)

You can route through your own OpenAI, Anthropic or Azure key, or an AWS Bedrock role. A human adds it once in the console as a connection. Your code still calls the gateway with a marketplace key. The provider bills you directly, and usage.cost is 0. A key bound to a connection can call only that connection's exact model ids, and it does not fall back to marketplace supply unless a routing policy allows it. See Bring your own key.

5. Keep sessions on one route and caches warm

A prompt cache saves money and time only when the next request reaches the same provider, or the same host, with the same prompt start. This is how the gateway treats sessions today:

  • The gateway keeps no session state. Your app stores the conversation and sends all of it on every turn.
  • Routing is deterministic. For the same model, key and provider object, every request tries the same route first. So the turns of a session stay on one provider without a session id.
  • A turn moves to the next route when the first route fails before it sends output (one failure is enough), when the route's capacity pool is full, or while its provider's 429 pause lasts (the provider's retry-after, 1 to 60 seconds). After two failures in a row, later turns skip the route for 30 seconds. A turn that moves can miss the cache once. On a follow-up turn, the gateway first waits up to 5 seconds for a full pool: see Waiting for the first route.
  • Send a session id, and the gateway gives it to each provider in the provider's own field: see Sessions and caching. Claude Code's x-claude-code-session-id header counts, with no change on your side.
  • For an open model, OpenRouter picks the host. With a session id (or the id that the gateway makes when you send none), OpenRouter keeps the conversation on one host from its first request, but it does not promise to. To force one host, pin it with upstream, as shown below. Remember that upstream turns off host fallback: if the pinned hosts fail, the request fails.

Rules for a high cache hit rate:

  1. 01
    Keep the start of the prompt byte-identical: the same system prompt, the same tool definitions in the same order, and earlier messages unchanged. Only append.
  2. 02
    Put content that changes (the time, ids, documents retrieved for this turn) after the stable part, at the end.
  3. 03
    Keep the model and the provider object the same for the whole session.
  4. 04
    For Claude, put cache_control on the block that ends the stable part (a system, tool or message block), or send the top-level cache_control field for automatic caching. Both surfaces keep the marks for Anthropic and OpenRouter: see Cache marks. On marketplace routes the cache lasts 5 minutes: ttl: "1h" is removed.
  5. 05
    OpenAI caches long repeated prompt starts by itself. Many open-model hosts do too. You set nothing for them.
  6. 06
    For an open model, pin one host for each session when cache hits matter more than host fallback.
  7. 07
    Measure. Cache reads are usage.prompt_tokens_details.cached_tokens on the OpenAI surface and usage.cache_read_input_tokens on the Anthropic surface. Cache writes are cache_write_tokens and cache_creation_input_tokens. Cache reads are billed at the lower cached_prompt price. Claude cache writes are billed at the higher cache_write price. Prices are in /api/models.json.

Python — the Anthropic SDK, with a cache mark on a long system prompt that does not change:

python
reply = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    system=[{
        "type": "text",
        "text": STABLE_INSTRUCTIONS,              # the same bytes on every turn
        "cache_control": {"type": "ephemeral"},   # cache everything up to here
    }],
    messages=history,                             # append only
)
print(reply.usage)  # cache_read_input_tokens, cache_creation_input_tokens and cost

The provider caches a prompt start only above a minimum length, so a short system prompt shows no cache reads.

For open models, give each session its own host. The same session always gets the same first host, and different sessions spread across hosts. The second host is used only when the first one fails:

python
import hashlib

# Hosts that run the model: its served_by entries in https://app.routerplus.com/api/models.json (tags as in step 4)
HOSTS = ["fireworks", "together", "deepinfra"]

def session_route(session_id: str) -> dict:
    i = int(hashlib.sha256(session_id.encode()).hexdigest(), 16) % len(HOSTS)
    return {"upstream": [HOSTS[i], HOSTS[(i + 1) % len(HOSTS)]]}

reply = client.chat.completions.create(
    model="moonshotai/kimi-k3",
    max_tokens=2048,
    messages=history,
    extra_body={"provider": session_route(session_id)},
)
print(reply.usage.prompt_tokens_details.cached_tokens)

Sessions and caching

Session ids

On POST /v1/chat/completions and POST /v1/messages, the session id is the first valid value among these inputs, in this order. Token counting, images and videos do not use it.

OrderInputKind
1session_idTop-level body field
2x-session-idHeader
3x-session-affinityHeader
4x-claude-code-session-idHeader that Claude Code sends by itself
  • A valid value has 1 to 256 printable ASCII characters. prompt_cache_key, a top-level body field, follows the same rule.
  • The gateway ignores an invalid value and uses the next input. An invalid value alone is never an error. The ignored input is named in x-tm-dropped-params as session_id (invalid), prompt_cache_key (invalid), header:x-session-id (invalid), header:x-session-affinity (invalid) or header:x-claude-code-session-id (invalid). With "require_parameters": true, that makes the request a 400, as any dropped field does.
  • Valid session inputs are never named in x-tm-dropped-params, and they never trigger require_parameters.
  • prompt_cache_key stays OpenAI's own field. It wins over the session id on OpenAI routes, and it is the fallback on OpenRouter.
  • When two valid inputs disagree, the order decides. There is no error.
  • Codex's session-id header is not read. It can matter only after the gateway adds POST /v1/responses, because current Codex calls only the Responses API.

What each route receives:

RouteFieldValue
openai (marketplace), and openai and azure (your connections)prompt_cache_keyYour prompt_cache_key, or else the session id. Nothing if you sent neither
openrouter (marketplace)session_idThe session id, or else your prompt_cache_key, or else an id that the gateway makes. If you sent prompt_cache_key, it goes too
anthropic and bedrockNothingAnthropic has no session input. Its cache matches the prompt start
  • Values go to the provider unchanged. A value longer than the provider accepts is replaced by tm- followed by 43 base64url characters: a hash of the value.
  • For OpenRouter only, when you send neither a session id nor prompt_cache_key, the gateway makes an id from your organization id, the model and the opening messages, up to the first user message. OpenRouter then keeps the conversation on one host from its first request.
  • The gateway keeps no session table. The id goes to the provider, and the provider routes on it.
  • prompt_cache_retention goes unchanged to openai and azure routes, both "in_memory" and "24h". On other routes it is dropped and named.

End-user ids on marketplace routes

On marketplace routes, the provider gets a hash in place of your end user's id, and the raw values do not go. The id is the first string among the body fields safety_identifier, user and metadata.user_id. The hash is tm- followed by 43 base64url characters, made from your organization id and that id. Without an id, it is made from your organization id alone. Each marketplace route gets exactly one field:

Marketplace routeField that carries the hashFields removed
openai, azuresafety_identifieruser
openrouterusersafety_identifier
anthropicmetadata.user_iduser, safety_identifier
every other route (OpenAI-compatible), dedicated capacity includedusersafety_identifier, metadata.user_id
  • A provider that blocks an id then blocks one user of one customer, not the whole marketplace account.
  • On your own connections, user and metadata pass unchanged. safety_identifier passes unchanged to openai and azure connections, and it is dropped and named on anthropic and bedrock connections.
  • Identity fields that a route uses are never named in x-tm-dropped-params.

Cache marks

The gateway never adds a cache mark that you did not send. Block and part marks (cache_control) are handled this way:

RequestRouteWhat happens to the marks
/v1/chat/completionsAnthropic direct or BedrockEach text part keeps its mark. A marked system part turns the system prompt into blocks. A tool message's last mark goes on its tool_result block. Other part keys are not forwarded
/v1/chat/completionsOpenRouterParts pass unchanged
/v1/chat/completionsOpenAI direct or AzurePart marks pass unchanged. OpenAI and Azure ignore them
/v1/messagesAnthropic direct or BedrockBlock marks pass unchanged
/v1/messagesOpenRouterMarks survive in system, user, assistant and tool content. Marks on tool definitions and on tool_use blocks are removed and named as tools[i].cache_control or messages[i].content[j].cache_control
/v1/messagesOpenAI direct or AzureMarks are removed and named

A top-level cache_control field (automatic caching), on either surface:

  • Anthropic direct and OpenRouter: sent unchanged.
  • Bedrock: turned into a mark on the last block that can carry one, counted back from the end of the messages. If no block can carry it, it is dropped and named as cache_control.
  • OpenAI direct and Azure: dropped and named as cache_control.
  • Any route: a field that Anthropic would refuse is dropped and named as cache_control. Anthropic refuses a field other than {"type": "ephemeral"} with an optional ttl of "5m" or "1h", a fifth mark (four marks exist already), a 1-hour field after a 5-minute mark, and a field whose TTL differs from the mark on the last block.

On marketplace routes, ttl: "1h" is removed from every mark, so the cache lasts 5 minutes. The drop is named as <path>.cache_control.ttl, for example system[0].cache_control.ttl, messages[2].content[0].cache_control.ttl, or cache_control.ttl for the top-level field. ttl: "5m" passes. Your own connections keep ttl: "1h".

Waiting for the first route

On a follow-up turn (an assistant or tool message after a user message), if the first route's capacity pool is full, the gateway waits for that pool for up to 5 seconds (an operator setting) before it moves the request to the next route. It never waits past the request's deadline. The x-tm-affinity-wait-ms response header gives the wait in milliseconds, only when the wait was more than 0.

6. Handle errors and retries

Every failure carries one stable class, error_type. Read it from error.metadata.error_type on the OpenAI surface, or error.error_type on the Anthropic surface. The x-tm-error-code header carries it too, but browsers cannot read that header. Never parse the message text. Log the x-request-id header with every failure.

error_typeHTTPRetryWhat to do
rate_limit429Yes, onceWait retry-after seconds. If x-tm-limit-kind is concurrency, send fewer requests at once.
email_not_verified403NoStop and tell a human to open the verify link in their email; the key starts working within seconds of the click.
insufficient_quota429NoStop and tell a human: add credits or raise a spend cap. On a cap, x-tm-cap-reset gives the reset time.
upstream_error5xxYes, once, after about 2 secondsThe gateway already tried every other provider it could. If the retry fails too, report it with the x-request-id.
upstream_error4xxNoThe provider refused the request itself. Read the message.
gateway_error503YesWait retry-after. The request did not reach a provider.
gateway_error504Yes, onceThe deadline passed while the gateway tried providers. An earlier attempt may be billed: check GET /v1/generation?id=<x-request-id> first.
gateway_error400NoSend one output limit, from 1 to 32,768.
gateway_error500NoReport it with the x-request-id.
model_unavailable404NoFix the model id, or fix your provider object (check it with POST /v1/route).
model_unavailable502Yes, onceThe connection to every provider failed.
invalid_request400NoFix the field that the message names.
context_overflow400NoShorten the input, or use a model with a larger context.
content_policy400NoChange the prompt. It is never sent to another provider.
auth401NoCheck the header, the key, and whether the key was disabled. A 401, 402 or 403 relayed from every provider is the marketplace's own account problem: report it with the x-request-id.
request_too_large413NoKeep the body under 10 MB (1 MB for a video request).
not_found404NoUse a route that exists. The list is in Errors.

Rules:

  1. 01
    Let the SDK retry. The official OpenAI and Anthropic SDKs retry 429 and 5xx responses twice by default, and they wait for retry-after. Keep that default. Do not add a second retry loop around it. The SDKs also retry an insufficient_quota response. That costs nothing, but your code must then stop.
  2. 02
    There is no idempotency key. A retried request is a new request. If it reaches a provider, it is billed again. A refusal by the gateway's own limits (x-tm-error-origin: gateway_admission) cost nothing.
  3. 03
    Never retry insufficient_quota, invalid_request, context_overflow, content_policy or auth in a loop. The request, or the account, must change first.
  4. 04
    An error in the middle of a stream is final for that answer. Send the whole turn again. Never join two partial answers.

Python — the OpenAI SDK, reading the class:

python
from openai import APIStatusError

try:
    reply = client.chat.completions.create(model=MODEL, max_tokens=1024, messages=messages)
except APIStatusError as e:
    error = e.response.json().get("error") or {}
    kind = (error.get("metadata") or {}).get("error_type") or e.response.headers.get("x-tm-error-code")
    request_id = e.response.headers.get("x-request-id")
    if kind == "insufficient_quota":
        notify_operator(request_id, error.get("message"))  # money or a cap: a retry cannot help
    raise

Details: Errors.

7. Limits, spend caps and your users

Default limits:

ScopeRequests per minuteTokens per minuteIn flight at once
Trial key201,000,0008
Console or identity API key3001,000,0008
Organization, all keys together3001,000,0008

Rules:

  1. 01
    Limit how many requests your code sends at once. By default the organization allows 8 in flight. A ninth gets 429 rate_limit with x-tm-limit-kind: concurrency. Use a semaphore or a fixed pool of workers. To run more at once, a human first raises the organization and key limits in /console/limits. A key's requests per minute cannot go above its ceiling (300 for a console key). Each marketplace provider pool also limits one organization to half of the pool's capacity. That refusal carries x-tm-limit-scope: pool.
  2. 02
    Read x-tm-remaining-rpm and x-tm-remaining-tpm on each response, and slow down before they reach zero. GET https://api.routerplus.com/v1/limits?model=<id> returns the limits that apply to your key.
  3. 03
    While a call runs, token limits count an estimate: one token for each four bytes of the request, rounded up, 65,536 for each image or document part, plus max_tokens. An image sent inline as base64 counts its 65,536 only, not its bytes. The real count replaces the estimate when the call ends. Small requests and a realistic max_tokens let more calls run at once.
  4. 04
    Spend caps are monthly, for each key and for the organization, on the UTC calendar month. A human sets them in the console. A cap refuses with 429 insufficient_quota and the x-tm-cap-reset header. Give each feature or customer its own key and cap, so a bug or a loop cannot spend everything.

Your users:

  • By default, one backend key serves all your users. The gateway does not know who your users are, so enforce per-user quotas in your app.
  • safety_identifier, user and metadata.user_id identify your end user to the provider. On marketplace routes the gateway sends only a hash of that id with your organization; on your own connections they go unchanged (see End-user ids on marketplace routes). They do not identify, limit or bill anyone on the gateway. Do not put emails or names in them.
  • For limits that the gateway enforces per user, use the identity API. Create an end_user principal for each user, issue a key for it, and set limits on the principal. Your backend keeps those keys. Never give a key to the user. The identity API takes a console session, not an API key, so a human with an owner or admin role sets it up. See Workspaces and principals.
  • To name your app in analytics, send the HTTP-Referer and X-Title headers. They carry no content.

Details: Rate limits and spend caps and Limits and capacity.

8. Track cost

  • Every billed response has usage.cost in USD. It is the full charge for the request, including provider attempts that failed before the answer. Store it with the x-request-id, the model and your own ids, such as the user and the feature.
  • In a stream, cost is in the usage chunk on the OpenAI surface, and in the last message_delta on the Anthropic surface.
  • An Anthropic-surface stream that fails in the middle of the answer has no usage in it. Read its cost from GET /v1/generation?id=<x-request-id>.
  • GET /v1/usage returns the balance and the last 20 attempts. For a key bound to its own workspace or principal, the balance fields are null. The console shows usage and logs for the whole organization.
  • BYOK calls show a cost of 0, because your provider bills you directly.
  • To check a charge, multiply the token counts by the prices in /api/models.json. The gateway rounds the final charge down.

Details: Pricing and billing and GET /v1/usage and /v1/generation.

9. Verify the integration

One cheap call shows the answer, the cost and the headers that the gateway adds:

bash
curl -sS -i https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"gpt-4o-mini","max_tokens":16,"messages":[{"role":"user","content":"Reply with OK"}]}'

Then confirm that the ledger saw it:

bash
curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usage

The integration is done when all of these are true:

  1. 01
    Every client uses the gateway base URL and reads the key from TM_API_KEY on the server.
  2. 02
    The key is not in the repository, in a client bundle or in a log.
  3. 03
    Every model id in the configuration appears in GET /v1/models.
  4. 04
    Every request sends max_tokens.
  5. 05
    One streamed and one non-streamed test call return usage.cost.
  6. 06
    The x-tm-dropped-params header is absent, or it names only fields that you accept losing.
  7. 07
    x-tm-provider, and x-tm-served-by when you pin a host, show the route that you expect. (x-tm-served-by uses OpenRouter's display name, not the tag.)
  8. 08
    The error handling switches on error_type. It stops on insufficient_quota, and it adds no retry loop on top of the SDK's own retries.
  9. 09
    The number of requests in flight stays within the organization's limit.
  10. 10
    The logs keep x-request-id and usage.cost for every call.
  11. 11
    The test calls appear in GET /v1/usage.

Leave instructions for the next agent

Add this block to the AGENTS.md or CLAUDE.md file of the repository that you integrated. Later coding sessions then keep the integration correct:

markdown
## LLM calls: RouterPlus gateway

- Every LLM call goes through the RouterPlus gateway. Full guide: https://app.routerplus.com/docs/agent-integration.md
- OpenAI SDK base URL: https://api.routerplus.com/v1. Anthropic SDK base URL: https://api.routerplus.com. Key: the TM_API_KEY environment variable, on the server only.
- Model ids live in configuration. Each one must exist in https://app.routerplus.com/api/models.json. Never invent or alias an id.
- Send max_tokens on every request (1 to 32768). Never send n above 1.
- Claude models: /v1/messages, with cache_control on the stable start of the prompt. GPT and author/model ids: /v1/chat/completions.
- Keep the start of the prompt byte-identical across turns. Only append.
- Send one id per conversation as the x-session-id header (or session_id in the body). Providers then keep its cache warm.
- Provider pins go in the request body "provider" object (only, ignore, order, allow_fallbacks, upstream, require_parameters). Check a pin with POST https://api.routerplus.com/v1/route.
- Errors: switch on error_type. rate_limit: wait retry-after, then retry once. insufficient_quota: stop and tell a human. Never retry invalid_request, context_overflow or content_policy.
- No idempotency key: a retry is a new, billed request. Add no retry loop on top of the SDK's own retries.
- At most 8 requests in flight per organization, unless the limit was raised.
- Log x-request-id and usage.cost for every call.
- Not available: the Responses API, embeddings, audio, files and batches.

Steps only a human can do

  • Complete browser signup through Clerk, including email verification, then buy credits in Billing. Signup and card verification alone do not add credit.
  • Fix billing when insufficient_quota appears: add credits, or raise a spend cap in the console.
  • Sign in to the console to create keys, set spend caps and change limits.
  • Add your own provider keys (BYOK), save routing policies, and set up workspaces and end-user principals.
  • Approve large test runs. Every call is billed.

Where next

Markdown source for agents: /docs/agent-integration.md · index at /llms.txt