Overview
What this is, both wire surfaces, and your first call.
RouterPlus is an LLM gateway: one API key and one prepaid balance in front of models from multiple providers. It speaks the two wire dialects your code already speaks — an OpenAI-compatible surface and an Anthropic-compatible surface — and every chat model in the catalog is callable from either one (image models answer on POST /v1/images/generations, video models on POST /v1/videos). The gateway translates requests, streams, and errors between dialects; your SDK never notices.
Pricing is pass-through: you pay the listed per-token rates, and the exact cost of every billed request is written into the response itself.
Two wire surfaces, one key
| Surface | Base URL | Completion endpoint | Works with |
|---|---|---|---|
| OpenAI-compatible | https://api.routerplus.com/v1 | POST /v1/chat/completions | OpenAI SDKs, Codex CLI, anything speaking the OpenAI wire format |
| Anthropic-compatible | https://api.routerplus.com | POST /v1/messages | Anthropic SDKs, Claude Code |
The same key authenticates on both, as Authorization: Bearer or x-api-key — see Authentication. The surface does not constrain the model: OpenAI-format code can call Claude models, Anthropic-format code can call GPT models. There is no lock between the dialect you speak and the model you get.
Your first call
Open Sign up in your browser, complete Clerk authentication and email verification, and save the first API key shown after sign-in. Add paid credits in Billing before your first call; signup does not add free credit. Existing customers can sign in and create another key at API keys.
We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match.
# Export the key saved from the browser.
export TM_API_KEY=tm_vk_...
# Call a model — streamed, billed, cost on the wire.
curl -N https://api.routerplus.com/v1/chat/completions \
-H "Authorization: Bearer $TM_API_KEY" \
-H "content-type: application/json" \
-d '{"model":"claude-haiku-4-5","stream":true,"max_tokens":60,"messages":[{"role":"user","content":"hello"}]}'The final usage chunk of the stream carries token counts and cost — the exact USD amount this request debited. When the balance is spent, visit Billing to add credits.
Or point your existing SDK at the gateway — the only changes are the base URL and the key:
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])
r = client.chat.completions.create(
model="claude-haiku-4-5", # yes — a Claude model over the OpenAI wire format
max_tokens=60,
messages=[{"role": "user", "content": "hello"}],
)
print(r.model_dump()["usage"]["cost"]) # exact USD debit for this requestimport os
import anthropic
client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])
msg = client.messages.create(
model="gpt-4o-mini", # and a GPT model over the Anthropic wire format
max_tokens=60,
messages=[{"role": "user", "content": "hello"}],
)Beyond the API
The site uses Clerk for browser signup and sign-in. Its pages run on the same catalog, routing, prices and ledger as the API:
- Playground — chat with any catalog model, compare up to three side by side, judge with a decision model, and render images and video. See Playground.
- Optimize — search for a cheaper model or mix that scores as well as yours on your own Braintrust evals, then deploy it. See Optimize.
- Endpoints — your organization's dedicated endpoints, then the models you deployed from Optimize, each a
tm/...model id your keys can call. - Console — usage, logs, API keys, provider connections, integrations, credits, spend caps and rate limits, one page per section.
What makes this gateway different
The exact cost of every request is on the wire
Every billed response carries usage.cost (USD) — in the final usage chunk on streams, in the response body otherwise. It is computed with the same integer micro-USD math the ledger settles with, and it covers the full request debit, including attempts that failed over before your answer started. You can recompute your bill from the wire at any time. Money math rounds in your favor: reservations round up, settlement rounds down.
GET /v1/usage lists your balance and recent requests with per-request cost; GET /v1/generation?id=<request id> audits every physical attempt behind one request. One documented limit: an Anthropic-surface stream that dies mid-answer has no legal wire slot for usage, so /v1/generation is the recomputation path there.
Failover with a commit boundary
Before any output has reached you, upstream failures fail over between deployments with zero backoff — invisibly, except for the x-tm-attempts header that counts physical dispatches. The moment the first real output reaches you, the request is committed to that provider forever. If the provider dies after that, you get one terminal error event inside the stream — never a silent restart on a different provider, never a mid-answer switch. A content-policy refusal is never rerouted to another provider, period: rerouting a refusal would be laundering it.
During long silences (slow reasoning models), the stream carries keep-alive frames every 15 seconds — an SSE comment on the OpenAI surface, a native ping event on the Anthropic surface — so proxies don't kill a healthy stream.
Strict where silence would cost you money
- An unknown model id is an honest 404 naming the id you sent — never a silent substitution with a model you didn't choose.
n>1is rejected with a 400 rather than billing you for a garbled single-choice response.- Content that cannot cross a dialect boundary (image and audio parts, in v1) is a typed 400 naming the field whenever translation is required — never silently dropped. So is a part the model does not take, such as an image to a model that reads text only. A parameter the other dialect has no mapping for (JSON mode on an Anthropic deployment, say) is a typed 400 too. Unknown top-level parameters are dropped on every route, and every drop is recorded: the
x-tm-dropped-paramsheader names them, and so does the request's audit row. - When every deployment serving a model is cooling down, you get a 503 with
retry-after— not a permanent-looking 404.
Content-free by design
Prompts and completions are never stored by the gateway. Metering, the usage endpoints, and telemetry record metadata only: token counts, timings, outcomes, cost. The optional HTTP-Referer and X-Title headers identify your app for analytics without exposing request content. One exception, stated where it applies: Optimize stores the eval cases and candidate answers of a search, for your organization, until you delete the search. See Data policy.
Uptime numbers that admit their sample size
Each model's page at https://app.routerplus.com/models/<id> shows per-provider uptime over the last 24 hours. For a provider we call directly, it shows a percentage only once that provider has at least 100 counted requests. Below that it shows n<100, because a percentage over a handful of requests is noise dressed as data. Buyer-caused failures (your 400s, your cancelled streams) never count against a provider's uptime. See The uptime figure is allowed to say nothing.
Endpoints at a glance
Gateway (https://api.routerplus.com) — all authenticated:
| Endpoint | What it does |
|---|---|
POST /v1/chat/completions | OpenAI-compatible completions, streaming and not |
POST /v1/messages | Anthropic-compatible messages, streaming and not |
POST /v1/messages/count_tokens | Anthropic token counting, unbilled; models with an Anthropic-dialect deployment only |
GET /v1/models | Models your key can call. OpenAI list shape by default; Anthropic shape when you send an anthropic-version header |
POST /v1/images/generations | OpenAI Images API: base64 in the body |
POST /v1/videos | OpenAI Videos API: a job, then GET /v1/videos/{id} and GET /v1/videos/{id}/content |
GET /v1/usage | Balance, credited/spent totals, recent requests with cost |
GET /v1/generation?id=<request id> | Per-attempt audit for one request: tokens, provenance, cost, outcome |
POST /v1/route | Explains how a request would route, without dispatching it. See Routing policies |
GET /v1/limits?model=<id> | The effective rate limits for your key. See Limits and capacity |
Site (https://app.routerplus.com):
| Endpoint | What it does |
|---|---|
/signup | Browser signup through Clerk, including email verification |
GET /api/models.json | Public catalog: ids, prices, context windows, who serves each model. No auth |
/login | Browser sign-in through Clerk |
/playground | Chat, decision, images and video with any catalog model |
/model-search | Optimize: model search on your Braintrust evals |
/endpoints | Your dedicated endpoints and your deployed tm/... models |
/console | Overview; then /console/usage, /console/logs, /console/keys, /console/byok, /console/integrations, /console/billing, /console/limits |
/llms.txt | Machine-readable docs index for agents |
/llms-full.txt | Every docs page as raw markdown, in one file |
Where next
error_type, what it means, and exactly what to do, including retry discipline for agents →Wire compatibility & billing contractThe conventions your bill is computed from; changing that page requires a recorded decision →Migrate in one promptRepoint the OpenAI SDK, Anthropic SDK, Claude Code, or Codex CLI with a single paste →Install (agent runbook)Hand it to your coding agent; it configures your client after browser signup →If a coding agent is doing the work, point it at https://app.routerplus.com/docs/agent-integration.md to build the gateway into a product, or at Install to wire a coding tool. https://app.routerplus.com/llms.txt is the index it fetches first; https://app.routerplus.com/llms-full.txt holds every page in one file.
Markdown source for agents: /docs/overview.md · index at /llms.txt