# Quickstart

Signup to a streamed, billed model call in five steps. You need a browser, `curl` and an email address.

The marketplace is one API key and one prepaid balance in front of two wire surfaces:

| Surface | Endpoint | Works with |
|---|---|---|
| OpenAI-compatible | `POST https://api.routerplus.com/v1/chat/completions` | OpenAI SDKs, anything OpenAI-shaped |
| Anthropic-compatible | `POST https://api.routerplus.com/v1/messages` | Anthropic SDKs, Claude Code |

Every chat model in the catalog is callable from **both** surfaces — the gateway translates requests, streams, and errors in either direction. Image models have their own route, [POST /v1/images/generations](/docs/api-images). Prices are pass-through, and every billed response carries `usage.cost` in USD, so you can recompute your bill from the wire.

> [!TIP]
> Setting up with a coding agent instead of by hand? Hand it the runbook on
> [Coding agents](/docs/install) — complete browser signup first, then let the agent configure your client.

## 1. Sign up in the browser and save your key

Open [Sign up](https://app.routerplus.com/signup) and complete Clerk authentication and email
verification. Your first verified sign-in shows your first `tm_vk_` API key. Save it then: only its hash is stored, so the
raw key cannot be shown again. Add paid credits in [Billing](https://app.routerplus.com/console/billing) before your first call. Signup does not add free credit.

We match your first $100 in credit purchases, dollar for dollar. The offer applies to eligible accounts; Billing shows your remaining match.

Already have an account? [Sign in](https://app.routerplus.com/login) with the same verified
email and create a key at [API keys](https://app.routerplus.com/console/keys). Your existing
balance and API keys are preserved.

Export the key:

```bash
export TM_API_KEY=tm_vk_...
```

Account creation requires the browser flow; the former `POST /v1/signup` route
has been removed. Once you have a key, all gateway calls below work from your
terminal or SDK. See [Authentication](/docs/authentication).

## 2. Pick a model

```bash
# Public catalog — no auth: ids, prices, context windows, provider retention
curl -s https://app.routerplus.com/api/models.json

# Authenticated — what your key can call, in your SDK's native list shape
curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/models
```

`GET /v1/models` returns the OpenAI list shape by default, and the Anthropic shape when you send an `anthropic-version` header. See [Models & catalog](/docs/models).

> [!NOTE]
> Model access is catalog-exact. An id that isn't listed returns a `404` with
> `error_type: model_unavailable`, echoing the id you asked for. The gateway never
> substitutes a "close enough" model — no silent aliasing, ever. If a request
> fails on the model id, the fix is the id, not a hidden routing preference.

## 3. First streamed call — OpenAI surface

A Claude model over the OpenAI wire format, to prove the translation is real:

```bash
curl -N https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-haiku-4-5",
    "stream": true,
    "max_tokens": 60,
    "messages": [{"role": "user", "content": "Say hello in five words."}]
  }'
```

The stream arrives in a fixed order: a role-priming delta, content deltas, a finish chunk, a **usage chunk**, then `data: [DONE]`. The usage chunk is your bill:

```json
{
  "prompt_tokens": 13,
  "completion_tokens": 9,
  "total_tokens": 22,
  "prompt_tokens_details": { "cached_tokens": 0, "cache_write_tokens": 0 },
  "completion_tokens_details": { "reasoning_tokens": 0 },
  "cost": 0.000058
}
```

`cost` is USD, computed with the same integer micro-USD math the ledger settles with, and it covers the full request — including any attempts that failed over before your answer started. Recompute it from the token counts and the public prices any time; [Pricing & billing](/docs/pricing) has the exact math.

Same call with the OpenAI Python SDK — the only changes from stock OpenAI are `base_url` and the key:

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])

stream = client.chat.completions.create(
    model="claude-haiku-4-5",
    max_tokens=60,
    stream=True,
    messages=[{"role": "user", "content": "Say hello in five words."}],
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")
    if chunk.usage:  # the final chunk before [DONE]
        print(f"\ncost: ${chunk.usage.cost}")
```

And TypeScript:

```typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.routerplus.com/v1",
  apiKey: process.env.TM_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "claude-haiku-4-5",
  max_tokens: 60,
  stream: true,
  messages: [{ role: "user", content: "Say hello in five words." }],
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
```

> [!TIP]
> Before dispatch, the gateway reserves the worst-case cost of the call —
> roughly (estimated input + `max_tokens`) at the model's prices — and settles
> down to observed usage afterward. On a small trial balance, set a sane
> `max_tokens` (it defaults to 4096, and 32,768 is the most a request may ask for): a huge value can make the reservation
> exceed your balance and return `insufficient_quota` before any provider is
> called. Details in [Rate limits & spend caps](/docs/limits).

## 4. Same key, Anthropic surface

A GPT model over the Anthropic wire format — the translation runs both ways:

```bash
curl -N https://api.routerplus.com/v1/messages \
  -H "x-api-key: $TM_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "stream": true,
    "max_tokens": 60,
    "messages": [{"role": "user", "content": "Say hello in five words."}]
  }'
```

The final `message_delta` event carries the usage: `input_tokens`, `output_tokens`, `cache_read_input_tokens`, `cache_creation_input_tokens`, and the same `cost` field in USD.

With the Anthropic Python SDK:

```python
import os
import anthropic

client = anthropic.Anthropic(
    base_url="https://api.routerplus.com",
    api_key=os.environ["TM_API_KEY"],
)

with client.messages.stream(
    model="gpt-4o-mini",
    max_tokens=60,
    messages=[{"role": "user", "content": "Say hello in five words."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="")
```

> [!NOTE]
> Both auth header styles work on both endpoints: `Authorization: Bearer` and
> `x-api-key`. Use whichever your SDK sends — no per-surface key juggling. See
> [Authentication](/docs/authentication).

## 5. What did it cost?

```bash
curl -s -H "Authorization: Bearer $TM_API_KEY" https://api.routerplus.com/v1/usage
```

```json
{
  "balance_usd": 4.999942,
  "credited_usd": 5.0,
  "spent_usd": 0.000058,
  "recent_attempts": [
    {
      "request_id": "5a2e…",
      "deployment": "anthropic",
      "model": "claude-haiku-4-5",
      "outcome": "completed",
      "usage_provenance": "observed",
      "input_tokens": 13,
      "output_tokens": 9,
      "cost_usd": 0.000058,
      "billing_source": "house",
      "at": "2026-09-04T10:14:03.201Z"
    }
  ]
}
```

The example is trimmed: each attempt also carries its price snapshot and inference cost, described on [GET /v1/usage](/docs/api-usage). `recent_attempts` lists your last 20 physical attempts. For the full audit of one request — every attempt including failovers, cache splits, reserved vs settled cost — use the `x-request-id` header from any response:

```bash
curl -s -H "Authorization: Bearer $TM_API_KEY" \
  "https://api.routerplus.com/v1/generation?id=REQUEST_ID"
```

Both endpoints are metadata-only: token counts, timings, outcomes, cost. Prompts and responses are never stored — see [Data policy](/docs/data-policy) and [GET /v1/usage & /v1/generation](/docs/api-usage).

## Deliberate strictness you'll notice

- Unknown model ids get a `404`, never a substitute ([Routing & failover](/docs/routing)).
- `n > 1` is rejected with a `400 invalid_request` rather than billing you for a garbled single-choice response.
- First keys shown after verified sign-in start at 20 requests/minute. Until your organization buys credit, all its keys together stay at 20 requests/minute; after the first purchase, keys created in the console get 300. Every organization also has token-per-minute and concurrency limits ([Rate limits & spend caps](/docs/limits)).
- Every error body carries a stable `error_type`, on both surfaces, in your SDK's native error shape ([Errors](/docs/errors)).

## Next steps

- [Agent integration guide](/docs/agent-integration) — the page an AI agent follows to build the gateway into your product
- [Coding agents](/docs/install) — point Claude Code or Codex at the gateway, then let an agent configure your client after browser signup
- [Migrate](/docs/migration) — repoint an existing codebase in one prompt
- [Streaming](/docs/streaming) — event order, keep-alives, and mid-stream failure semantics
- [Pricing & billing](/docs/pricing) — the exact math behind `usage.cost`
- [Wire compatibility](/docs/compat) — the contract your bill is computed from
