Console
Get started/Quickstart

Quickstart

Signup to streamed call in five steps.

/llms.txt

Signup to a streamed, billed model call in five steps. You need a browser, curl and an email address.

The marketplace is one API key and one prepaid balance in front of two wire surfaces:

SurfaceEndpointWorks with
OpenAI-compatiblePOST https://api.routerplus.com/v1/chat/completionsOpenAI SDKs, anything OpenAI-shaped
Anthropic-compatiblePOST https://api.routerplus.com/v1/messagesAnthropic SDKs, Claude Code

Every chat model in the catalog is callable from both surfaces — the gateway translates requests, streams, and errors in either direction. Image models have their own route, POST /v1/images/generations. Prices are pass-through, and every billed response carries usage.cost in USD, so you can recompute your bill from the wire.

Tip

Setting up with a coding agent instead of by hand? Hand it the runbook on Coding agents — complete browser signup first, then let the agent configure your client.

1. Sign up in the browser and save your key

Open Sign up and complete Clerk authentication and email verification. Your first verified sign-in shows your first tm_vk_ API key. Save it then: only its hash is stored, so the raw key cannot be shown again. Add paid credits in Billing before your first call. Signup does not add free credit.

We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match.

Already have an account? Sign in with the same verified email and create a key at API keys. Your existing balance and API keys are preserved.

Export the key:

bash
export TM_API_KEY=tm_vk_...

Account creation requires the browser flow; the former POST /v1/signup route has been removed. Once you have a key, all gateway calls below work from your terminal or SDK. See Authentication.

2. Pick a model

bash
# Public catalog — no auth: ids, prices, context windows, provider retention
curl -s https://app.routerplus.com/api/models.json

# Authenticated — what your key can call, in your SDK's native list shape
curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/models

GET /v1/models returns the OpenAI list shape by default, and the Anthropic shape when you send an anthropic-version header. See Models & catalog.

Note

Model access is catalog-exact. An id that isn't listed returns a 404 with error_type: model_unavailable, echoing the id you asked for. The gateway never substitutes a "close enough" model — no silent aliasing, ever. If a request fails on the model id, the fix is the id, not a hidden routing preference.

3. First streamed call — OpenAI surface

A Claude model over the OpenAI wire format, to prove the translation is real:

bash
curl -N https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-haiku-4-5",
    "stream": true,
    "max_tokens": 60,
    "messages": [{"role": "user", "content": "Say hello in five words."}]
  }'

The stream arrives in a fixed order: a role-priming delta, content deltas, a finish chunk, a usage chunk, then data: [DONE]. The usage chunk is your bill:

json
{
  "prompt_tokens": 13,
  "completion_tokens": 9,
  "total_tokens": 22,
  "prompt_tokens_details": { "cached_tokens": 0, "cache_write_tokens": 0 },
  "completion_tokens_details": { "reasoning_tokens": 0 },
  "cost": 0.000058
}

cost is USD, computed with the same integer micro-USD math the ledger settles with, and it covers the full request — including any attempts that failed over before your answer started. Recompute it from the token counts and the public prices any time; Pricing & billing has the exact math.

Same call with the OpenAI Python SDK — the only changes from stock OpenAI are base_url and the key:

python
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])

stream = client.chat.completions.create(
    model="claude-haiku-4-5",
    max_tokens=60,
    stream=True,
    messages=[{"role": "user", "content": "Say hello in five words."}],
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")
    if chunk.usage:  # the final chunk before [DONE]
        print(f"\ncost: ${chunk.usage.cost}")
typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.routerplus.com/v1",
  apiKey: process.env.TM_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "claude-haiku-4-5",
  max_tokens: 60,
  stream: true,
  messages: [{ role: "user", content: "Say hello in five words." }],
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
Tip

Before dispatch, the gateway reserves the worst-case cost of the call — roughly (estimated input + max_tokens) at the model's prices — and settles down to observed usage afterward. On a small trial balance, set a sane max_tokens (it defaults to 4096, and 32,768 is the most a request may ask for): a huge value can make the reservation exceed your balance and return insufficient_quota before any provider is called. Details in Rate limits & spend caps.

4. Same key, Anthropic surface

A GPT model over the Anthropic wire format — the translation runs both ways:

bash
curl -N https://api.routerplus.com/v1/messages \
  -H "x-api-key: $TM_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "stream": true,
    "max_tokens": 60,
    "messages": [{"role": "user", "content": "Say hello in five words."}]
  }'

The final message_delta event carries the usage: input_tokens, output_tokens, cache_read_input_tokens, cache_creation_input_tokens, and the same cost field in USD.

With the Anthropic Python SDK:

python
import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://api.routerplus.com",
    api_key=os.environ["TM_API_KEY"],
)

with client.messages.stream(
    model="gpt-4o-mini",
    max_tokens=60,
    messages=[{"role": "user", "content": "Say hello in five words."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="")
Note

Both auth header styles work on both endpoints: Authorization: Bearer and x-api-key. Use whichever your SDK sends — no per-surface key juggling. See Authentication.

5. What did it cost?

bash
curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usage
json
{
  "balance_usd": 4.999942,
  "credited_usd": 5.0,
  "spent_usd": 0.000058,
  "recent_attempts": [
    {
      "request_id": "5a2e…",
      "deployment": "anthropic",
      "model": "claude-haiku-4-5",
      "outcome": "completed",
      "usage_provenance": "observed",
      "input_tokens": 13,
      "output_tokens": 9,
      "cost_usd": 0.000058,
      "billing_source": "house",
      "at": "2026-09-04T10:14:03.201Z"
    }
  ]
}

The example is trimmed: each attempt also carries its price snapshot and inference cost, described on GET /v1/usage. recent_attempts lists your last 20 physical attempts. For the full audit of one request — every attempt including failovers, cache splits, reserved vs settled cost — use the x-request-id header from any response:

bash
curl -s -H "Authorization: Bearer $TM_API_KEY" \
  "https://api.routerplus.com/v1/generation?id=REQUEST_ID"

Both endpoints are metadata-only: token counts, timings, outcomes, cost. Prompts and responses are never stored — see Data policy and GET /v1/usage & /v1/generation.

Deliberate strictness you'll notice

  • Unknown model ids get a 404, never a substitute (Routing & failover).
  • n > 1 is rejected with a 400 invalid_request rather than billing you for a garbled single-choice response.
  • First keys shown after verified sign-in start at 20 requests/minute. Until your organization buys credit, all its keys together stay at 20 requests/minute; after the first purchase, keys created in the console get 300. Every organization also has token-per-minute and concurrency limits (Rate limits & spend caps).
  • Every error body carries a stable error_type, on both surfaces, in your SDK's native error shape (Errors).

Next steps

Markdown source for agents: /docs/quickstart.md · index at /llms.txt