Console
Get started/Overview

Overview

What this is, both wire surfaces, and your first call.

/llms.txt

RouterPlus is an LLM gateway: one API key and one prepaid balance in front of models from multiple providers. It speaks the two wire dialects your code already speaks — an OpenAI-compatible surface and an Anthropic-compatible surface — and every chat model in the catalog is callable from either one (image models answer on POST /v1/images/generations, video models on POST /v1/videos). The gateway translates requests, streams, and errors between dialects; your SDK never notices.

Pricing is pass-through: you pay the listed per-token rates, and the exact cost of every billed request is written into the response itself.

Two wire surfaces, one key

SurfaceBase URLCompletion endpointWorks with
OpenAI-compatiblehttps://api.routerplus.com/v1POST /v1/chat/completionsOpenAI SDKs, Codex CLI, anything speaking the OpenAI wire format
Anthropic-compatiblehttps://api.routerplus.comPOST /v1/messagesAnthropic SDKs, Claude Code

The same key authenticates on both, as Authorization: Bearer or x-api-key — see Authentication. The surface does not constrain the model: OpenAI-format code can call Claude models, Anthropic-format code can call GPT models. There is no lock between the dialect you speak and the model you get.

Your first call

Open Sign up in your browser, complete Clerk authentication and email verification, and save the first API key shown after sign-in. Add paid credits in Billing before your first call; signup does not add free credit. Existing customers can sign in and create another key at API keys.

We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match.

bash
# Export the key saved from the browser.
export TM_API_KEY=tm_vk_...

# Call a model — streamed, billed, cost on the wire.
curl -N https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"claude-haiku-4-5","stream":true,"max_tokens":60,"messages":[{"role":"user","content":"hello"}]}'

The final usage chunk of the stream carries token counts and cost — the exact USD amount this request debited. When the balance is spent, visit Billing to add credits.

Or point your existing SDK at the gateway — the only changes are the base URL and the key:

python
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])
r = client.chat.completions.create(
    model="claude-haiku-4-5",   # yes — a Claude model over the OpenAI wire format
    max_tokens=60,
    messages=[{"role": "user", "content": "hello"}],
)
print(r.model_dump()["usage"]["cost"])  # exact USD debit for this request
python
import os
import anthropic

client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])
msg = client.messages.create(
    model="gpt-4o-mini",        # and a GPT model over the Anthropic wire format
    max_tokens=60,
    messages=[{"role": "user", "content": "hello"}],
)

Beyond the API

The site uses Clerk for browser signup and sign-in. Its pages run on the same catalog, routing, prices and ledger as the API:

  • Playground — chat with any catalog model, compare up to three side by side, judge with a decision model, and render images and video. See Playground.
  • Optimize — search for a cheaper model or mix that scores as well as yours on your own Braintrust evals, then deploy it. See Optimize.
  • Endpoints — your organization's dedicated endpoints, then the models you deployed from Optimize, each a tm/... model id your keys can call.
  • Console — usage, logs, API keys, provider connections, integrations, credits, spend caps and rate limits, one page per section.

What makes this gateway different

The exact cost of every request is on the wire

Every billed response carries usage.cost (USD) — in the final usage chunk on streams, in the response body otherwise. It is computed with the same integer micro-USD math the ledger settles with, and it covers the full request debit, including attempts that failed over before your answer started. You can recompute your bill from the wire at any time. Money math rounds in your favor: reservations round up, settlement rounds down.

GET /v1/usage lists your balance and recent requests with per-request cost; GET /v1/generation?id=<request id> audits every physical attempt behind one request. One documented limit: an Anthropic-surface stream that dies mid-answer has no legal wire slot for usage, so /v1/generation is the recomputation path there.

Failover with a commit boundary

Before any output has reached you, upstream failures fail over between deployments with zero backoff — invisibly, except for the x-tm-attempts header that counts physical dispatches. The moment the first real output reaches you, the request is committed to that provider forever. If the provider dies after that, you get one terminal error event inside the stream — never a silent restart on a different provider, never a mid-answer switch. A content-policy refusal is never rerouted to another provider, period: rerouting a refusal would be laundering it.

During long silences (slow reasoning models), the stream carries keep-alive frames every 15 seconds — an SSE comment on the OpenAI surface, a native ping event on the Anthropic surface — so proxies don't kill a healthy stream.

Strict where silence would cost you money

  • An unknown model id is an honest 404 naming the id you sent — never a silent substitution with a model you didn't choose.
  • n>1 is rejected with a 400 rather than billing you for a garbled single-choice response.
  • Content that cannot cross a dialect boundary (image and audio parts, in v1) is a typed 400 naming the field whenever translation is required — never silently dropped. So is a part the model does not take, such as an image to a model that reads text only. A parameter the other dialect has no mapping for (JSON mode on an Anthropic deployment, say) is a typed 400 too. Unknown top-level parameters are dropped on every route, and every drop is recorded: the x-tm-dropped-params header names them, and so does the request's audit row.
  • When every deployment serving a model is cooling down, you get a 503 with retry-after — not a permanent-looking 404.

Content-free by design

Prompts and completions are never stored by the gateway. Metering, the usage endpoints, and telemetry record metadata only: token counts, timings, outcomes, cost. The optional HTTP-Referer and X-Title headers identify your app for analytics without exposing request content. One exception, stated where it applies: Optimize stores the eval cases and candidate answers of a search, for your organization, until you delete the search. See Data policy.

Uptime numbers that admit their sample size

Each model's page at https://app.routerplus.com/models/<id> shows per-provider uptime over the last 24 hours. For a provider we call directly, it shows a percentage only once that provider has at least 100 counted requests. Below that it shows n<100, because a percentage over a handful of requests is noise dressed as data. Buyer-caused failures (your 400s, your cancelled streams) never count against a provider's uptime. See The uptime figure is allowed to say nothing.

Endpoints at a glance

Gateway (https://api.routerplus.com) — all authenticated:

EndpointWhat it does
POST /v1/chat/completionsOpenAI-compatible completions, streaming and not
POST /v1/messagesAnthropic-compatible messages, streaming and not
POST /v1/messages/count_tokensAnthropic token counting, unbilled; models with an Anthropic-dialect deployment only
GET /v1/modelsModels your key can call. OpenAI list shape by default; Anthropic shape when you send an anthropic-version header
POST /v1/images/generationsOpenAI Images API: base64 in the body
POST /v1/videosOpenAI Videos API: a job, then GET /v1/videos/{id} and GET /v1/videos/{id}/content
GET /v1/usageBalance, credited/spent totals, recent requests with cost
GET /v1/generation?id=<request id>Per-attempt audit for one request: tokens, provenance, cost, outcome
POST /v1/routeExplains how a request would route, without dispatching it. See Routing policies
GET /v1/limits?model=<id>The effective rate limits for your key. See Limits and capacity

Site (https://app.routerplus.com):

EndpointWhat it does
/signupBrowser signup through Clerk, including email verification
GET /api/models.jsonPublic catalog: ids, prices, context windows, who serves each model. No auth
/loginBrowser sign-in through Clerk
/playgroundChat, decision, images and video with any catalog model
/model-searchOptimize: model search on your Braintrust evals
/endpointsYour dedicated endpoints and your deployed tm/... models
/consoleOverview; then /console/usage, /console/logs, /console/keys, /console/byok, /console/integrations, /console/billing, /console/limits
/llms.txtMachine-readable docs index for agents
/llms-full.txtEvery docs page as raw markdown, in one file

Where next

Tip

If a coding agent is doing the work, point it at https://app.routerplus.com/docs/agent-integration.md to build the gateway into a product, or at Install to wire a coding tool. https://app.routerplus.com/llms.txt is the index it fetches first; https://app.routerplus.com/llms-full.txt holds every page in one file.

Markdown source for agents: /docs/overview.md · index at /llms.txt