Quickstart
Signup to streamed call in five steps.
Signup to a streamed, billed model call in five steps. You need a browser, curl and an email address.
The marketplace is one API key and one prepaid balance in front of two wire surfaces:
| Surface | Endpoint | Works with |
|---|---|---|
| OpenAI-compatible | POST https://api.routerplus.com/v1/chat/completions | OpenAI SDKs, anything OpenAI-shaped |
| Anthropic-compatible | POST https://api.routerplus.com/v1/messages | Anthropic SDKs, Claude Code |
Every chat model in the catalog is callable from both surfaces — the gateway translates requests, streams, and errors in either direction. Image models have their own route, POST /v1/images/generations. Prices are pass-through, and every billed response carries usage.cost in USD, so you can recompute your bill from the wire.
Setting up with a coding agent instead of by hand? Hand it the runbook on Coding agents — complete browser signup first, then let the agent configure your client.
1. Sign up in the browser and save your key
Open Sign up and complete Clerk authentication and email verification. Your first verified sign-in shows your first tm_vk_ API key. Save it then: only its hash is stored, so the raw key cannot be shown again. Add paid credits in Billing before your first call. Signup does not add free credit.
We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match.
Already have an account? Sign in with the same verified email and create a key at API keys. Your existing balance and API keys are preserved.
Export the key:
export TM_API_KEY=tm_vk_...Account creation requires the browser flow; the former POST /v1/signup route has been removed. Once you have a key, all gateway calls below work from your terminal or SDK. See Authentication.
2. Pick a model
# Public catalog — no auth: ids, prices, context windows, provider retention
curl -s https://app.routerplus.com/api/models.json
# Authenticated — what your key can call, in your SDK's native list shape
curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/modelsGET /v1/models returns the OpenAI list shape by default, and the Anthropic shape when you send an anthropic-version header. See Models & catalog.
Model access is catalog-exact. An id that isn't listed returns a 404 with error_type: model_unavailable, echoing the id you asked for. The gateway never substitutes a "close enough" model — no silent aliasing, ever. If a request fails on the model id, the fix is the id, not a hidden routing preference.
3. First streamed call — OpenAI surface
A Claude model over the OpenAI wire format, to prove the translation is real:
curl -N https://api.routerplus.com/v1/chat/completions \
-H "Authorization: Bearer $TM_API_KEY" \
-H "content-type: application/json" \
-d '{
"model": "claude-haiku-4-5",
"stream": true,
"max_tokens": 60,
"messages": [{"role": "user", "content": "Say hello in five words."}]
}'The stream arrives in a fixed order: a role-priming delta, content deltas, a finish chunk, a usage chunk, then data: [DONE]. The usage chunk is your bill:
{
"prompt_tokens": 13,
"completion_tokens": 9,
"total_tokens": 22,
"prompt_tokens_details": { "cached_tokens": 0, "cache_write_tokens": 0 },
"completion_tokens_details": { "reasoning_tokens": 0 },
"cost": 0.000058
}cost is USD, computed with the same integer micro-USD math the ledger settles with, and it covers the full request — including any attempts that failed over before your answer started. Recompute it from the token counts and the public prices any time; Pricing & billing has the exact math.
Same call with the OpenAI Python SDK — the only changes from stock OpenAI are base_url and the key:
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])
stream = client.chat.completions.create(
model="claude-haiku-4-5",
max_tokens=60,
stream=True,
messages=[{"role": "user", "content": "Say hello in five words."}],
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
if chunk.usage: # the final chunk before [DONE]
print(f"\ncost: ${chunk.usage.cost}")import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.routerplus.com/v1",
apiKey: process.env.TM_API_KEY,
});
const stream = await client.chat.completions.create({
model: "claude-haiku-4-5",
max_tokens: 60,
stream: true,
messages: [{ role: "user", content: "Say hello in five words." }],
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Before dispatch, the gateway reserves the worst-case cost of the call — roughly (estimated input + max_tokens) at the model's prices — and settles down to observed usage afterward. On a small trial balance, set a sane max_tokens (it defaults to 4096, and 32,768 is the most a request may ask for): a huge value can make the reservation exceed your balance and return insufficient_quota before any provider is called. Details in Rate limits & spend caps.
4. Same key, Anthropic surface
A GPT model over the Anthropic wire format — the translation runs both ways:
curl -N https://api.routerplus.com/v1/messages \
-H "x-api-key: $TM_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "gpt-4o-mini",
"stream": true,
"max_tokens": 60,
"messages": [{"role": "user", "content": "Say hello in five words."}]
}'The final message_delta event carries the usage: input_tokens, output_tokens, cache_read_input_tokens, cache_creation_input_tokens, and the same cost field in USD.
With the Anthropic Python SDK:
import os
import anthropic
client = anthropic.Anthropic(
base_url="https://api.routerplus.com",
api_key=os.environ["TM_API_KEY"],
)
with client.messages.stream(
model="gpt-4o-mini",
max_tokens=60,
messages=[{"role": "user", "content": "Say hello in five words."}],
) as stream:
for text in stream.text_stream:
print(text, end="")Both auth header styles work on both endpoints: Authorization: Bearer and x-api-key. Use whichever your SDK sends — no per-surface key juggling. See Authentication.
5. What did it cost?
curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usage{
"balance_usd": 4.999942,
"credited_usd": 5.0,
"spent_usd": 0.000058,
"recent_attempts": [
{
"request_id": "5a2e…",
"deployment": "anthropic",
"model": "claude-haiku-4-5",
"outcome": "completed",
"usage_provenance": "observed",
"input_tokens": 13,
"output_tokens": 9,
"cost_usd": 0.000058,
"billing_source": "house",
"at": "2026-09-04T10:14:03.201Z"
}
]
}The example is trimmed: each attempt also carries its price snapshot and inference cost, described on GET /v1/usage. recent_attempts lists your last 20 physical attempts. For the full audit of one request — every attempt including failovers, cache splits, reserved vs settled cost — use the x-request-id header from any response:
curl -s -H "Authorization: Bearer $TM_API_KEY" \
"https://api.routerplus.com/v1/generation?id=REQUEST_ID"Both endpoints are metadata-only: token counts, timings, outcomes, cost. Prompts and responses are never stored — see Data policy and GET /v1/usage & /v1/generation.
Deliberate strictness you'll notice
- Unknown model ids get a
404, never a substitute (Routing & failover). n > 1is rejected with a400 invalid_requestrather than billing you for a garbled single-choice response.- First keys shown after verified sign-in start at 20 requests/minute. Until your organization buys credit, all its keys together stay at 20 requests/minute; after the first purchase, keys created in the console get 300. Every organization also has token-per-minute and concurrency limits (Rate limits & spend caps).
- Every error body carries a stable
error_type, on both surfaces, in your SDK's native error shape (Errors).
Next steps
usage.cost →Wire compatibilityThe contract your bill is computed from →Markdown source for agents: /docs/quickstart.md · index at /llms.txt