# POST /v1/messages

The Anthropic-compatible surface. Point the Anthropic SDK (or Claude Code) at `https://api.routerplus.com` with your marketplace key and it works — including against GPT models. Every chat model in the catalog is callable from this endpoint regardless of which provider serves it; image models answer on [POST /v1/images/generations](/docs/api-images) only, and video models on [POST /v1/videos](/docs/api-videos) only; the gateway translates requests, streams, and errors between dialects. The OpenAI-native twin is [POST /v1/chat/completions](/docs/api-chat-completions).

```
POST https://api.routerplus.com/v1/messages
```

Query strings are tolerated — clients that append them (Claude Code sends `/v1/messages?beta=true`) route normally.

## Headers

| Header | Required | Behavior |
|---|---|---|
| `x-api-key` | one of these two | Your marketplace key, `tm_vk_...` — the Anthropic-native header. |
| `Authorization` | | `Bearer tm_vk_...` also accepted, same key. |
| `content-type` | yes | `application/json`. |
| `anthropic-version` | no | Accepted for SDK compatibility; the gateway speaks `anthropic-version: 2023-06-01` to Anthropic upstreams regardless of what you send. |
| `anthropic-beta` | no | Forwarded when an Anthropic-dialect provider serves the request, for the values that do not change how a token is billed: `prompt-caching`, `token-efficient-tools`, `fine-grained-tool-streaming`, `interleaved-thinking`, `claude-code`, `oauth` and `computer-use` prefixes. Any other value is dropped and recorded in `x-tm-dropped-params` as `header:anthropic-beta:<value>`. No effect on other routes. |

A missing or invalid key returns 401 in the Anthropic error shape with `error_type: "auth"`. Optionally send `HTTP-Referer` and `X-Title` for content-free app attribution.

## Request

JSON body, capped at 10 MB (413 `request_too_large` above that).

| Parameter | Type | Behavior |
|---|---|---|
| `model` | string, required | A catalog id — any chat model, not just Claude. Unknown ids are an honest 404; an image model is a 400 naming `/v1/images/generations`, a video model a 400 naming `/v1/videos`. A deployed endpoint id (`tm/<name>-v<n>`) is served on `POST /v1/chat/completions` only and is a 400 here. |
| `max_tokens` | number | Output cap, 1 to 32,768. **When you omit it, the gateway sets `max_tokens: 4096`** on the request it forwards, whichever provider serves it. A value above 32,768 is a 400. |
| `messages` | array, required | Roles `user` and `assistant`. Content: a string, or blocks — `text`, `tool_use` (assistant), `tool_result` (user), `thinking` / `redacted_thinking` (assistant). `image` and `document` blocks are for a model that takes them (`architecture.input_modalities` in [`GET /v1/models`](/docs/api-models)): a block the model does not take is a typed 400 naming it, also inside a `tool_result`. `image` blocks are a typed 400 when translation to an OpenAI-dialect provider is needed; on Anthropic-dialect routes the messages pass through as sent. |
| `system` | string \| array | A string or text blocks. |
| `stream` | boolean | SSE streaming; see below. |
| `temperature`, `top_p` | number | Passed through. |
| `stop_sequences` | string[] | Becomes OpenAI `stop` on cross-dialect routes. |
| `tools` | array | `{name, description?, input_schema}`. Anthropic server tools (computer use, web search, bash, text editor) pass through on Anthropic-dialect routes and are a typed 400 when an OpenAI-dialect provider would serve the request. |
| `tool_choice` | object | `{"type":"auto"}`, `{"type":"any"}`, `{"type":"none"}`, or `{"type":"tool","name":"..."}`. `disable_parallel_tool_use: true` becomes `parallel_tool_calls: false` on OpenAI-dialect routes. |
| `thinking`, `output_config` | object | Pass through on Anthropic-dialect routes; a typed 400 when translation to an OpenAI-dialect provider is needed — the gateway does not map them yet and will not drop them silently. |
| `top_k`, `metadata` | | Pass through on Anthropic-dialect routes; dropped and recorded on OpenAI-dialect routes. On a house route, `metadata.user_id` becomes a hash with your organization; see [Sessions and prompt caching](/docs/compat#sessions-and-prompt-caching). |
| `cache_control` | object | Automatic prompt caching, on the request or on blocks. Passes to Anthropic and OpenRouter (on Bedrock the request-level field becomes a mark on the last block); dropped and recorded on OpenAI and Azure. A request-level field that Anthropic would refuse is dropped and recorded on every route (see [Wire compatibility](/docs/compat)). On a house route, `ttl: "1h"` is removed (5 minutes) and recorded. |
| `session_id` | string | A session id, also accepted as the header `x-session-id`, `x-session-affinity` or `x-claude-code-session-id`. It reaches OpenRouter as `session_id` and OpenAI or Azure as `prompt_cache_key`; Anthropic has no session input. See [Sessions and prompt caching](/docs/compat#sessions-and-prompt-caching). |
| `mcp_servers`, `container` | | A typed 400 on every route: stripping them would change who runs what. |
| `provider` | object | Routing controls read by the gateway and never forwarded: `require_parameters`, `order`, `only`, `ignore`, `allow_fallbacks`, `upstream` and the rest — see [Routing policies](/docs/routing-policies). An unknown control is a 400. |

Any other top-level key is dropped and recorded on every route (the `x-tm-dropped-params` header names it), never mutated. A top-level `api_key`, `base_url` or `connection_id` is a 400. Unsupported **content** is never dropped: a block the model does not take, or one a translation cannot carry, is a typed 400 naming the exact field, because dropped content would still be billed upstream. A typed 400 from translation is decided per deployment: when another deployment of the model speaks the Anthropic dialect, the request goes there instead — see [Wire compatibility](/docs/compat).

> [!NOTE]
> Transcripts with `thinking` blocks round-trip safely. When a model served over the OpenAI dialect streams reasoning, this surface encodes it as `thinking` blocks — and when you echo that transcript back and the request routes to an OpenAI-dialect provider, assistant `thinking`/`redacted_thinking` blocks are dropped rather than rejected. The route that produced your transcript will never 400 on it.

## Streamed response

With `stream: true`, the real Anthropic Messages event sequence — the same normalized encoding whichever provider serves the request. On a translated stream the message `id` is `msg_<request-id>`, joinable with the `x-request-id` header and `/v1/generation?id=`.

```text
event: message_start
data: {"type":"message_start","message":{"id":"msg_1c9a7b2e-...","type":"message","role":"assistant","model":"claude-sonnet-4-5","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":0,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"input_tokens":12,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":5,"cost":0.000111}}

event: message_stop
data: {"type":"message_stop"}
```

Block types you may see: `text` (`text_delta`), `thinking` (`thinking_delta`, plus `signature_delta` for multi-turn thinking with tools), `redacted_thinking`, and `tool_use` (`input_json_delta` fragments). Streamed `stop_reason` is one of `end_turn`, `max_tokens`, `tool_use`, `refusal`.

Usage fields, per the [billing contract](/docs/compat): `input_tokens` **excludes** cache reads and writes; `cache_read_input_tokens` and `cache_creation_input_tokens` carry the cache splits; `usage.cost` (USD, billed keys only) in the final `message_delta` is the full request debit, failed-over attempts included, computed with the exact math the ledger settles with.

> [!NOTE]
> Treat the final `message_delta`'s usage as authoritative. On same-dialect routes (a Claude model on this surface) the stream is relayed record-for-record, so `message_start` carries the provider's real token counts and its own message id. On cross-dialect routes (an OpenAI-dialect provider behind this surface) the gateway re-encodes the stream and `message_start`'s counts are zeros — the final `message_delta` always carries the real numbers on every route.

During post-start silences of 15 s or more the gateway emits native `ping` events (`{"type":"ping"}`) so idle-but-healthy streams survive proxies. If the serving provider dies **after** output started, you get one terminal `error` event inside the 200 stream — carrying the Anthropic-native `type` plus a stable `error_type` — and nothing after it; the answer is never silently restarted on another provider. The same event ends a stream that reaches the request deadline (450 s from dispatch, or a routing policy's `timeout_ms`). One documented limit: a billed stream that dies mid-answer has no legal Anthropic wire slot for usage outside `message_delta`, so recompute that request's cost from `/v1/generation?id=`.

## Non-streamed response

```json
{
  "id": "msg_1c9a7b2e-...",
  "type": "message",
  "role": "assistant",
  "model": "gpt-4o-mini",
  "content": [{ "type": "text", "text": "Hello!" }],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 12,
    "output_tokens": 5,
    "cache_read_input_tokens": 0,
    "cache_creation_input_tokens": 0,
    "cost": 0.000004
  }
}
```

`stop_reason` on translated responses is `end_turn`, `max_tokens`, `tool_use`, or `refusal`; when an Anthropic-dialect provider serves the request the body passes through with its native values (including `stop_sequence`) plus the injected `cost`.

## Errors

Anthropic error shape with the native `type` string your SDK switches on (`authentication_error`, `rate_limit_error`, `billing_error`, `not_found_error`, ...) plus a stable `error_type` carrying the canonical class:

```json
{
  "type": "error",
  "error": {
    "type": "not_found_error",
    "message": "model \"claude-9\" is not in the catalog; GET /v1/models lists what this key can serve",
    "error_type": "model_unavailable"
  },
  "request_id": "..."
}
```

Out-of-credits and spend-cap breaches are 429 with native type `billing_error` and `error_type: "insufficient_quota"` (cap breaches include the exact UTC reset time and an `x-tm-cap-reset` header). A key whose signup email is not verified yet gets 403 with native type `permission_error` and `error_type: "email_not_verified"`. Until an organization has added credit, its keys and the organization as a whole run at 20 RPM — 429 `rate_limit` with `retry-after`. A provider's own error body is never relayed; when a provider rejected the request, the class says so and `x-tm-upstream-status` carries its status. Full table: [Errors](/docs/errors). Every response carries `x-request-id`; served requests add `x-tm-provider` and `x-tm-attempts`, failures add `x-tm-error-code` (and `x-tm-error-origin` when a limit, a provider or our infrastructure refused the request), and `x-tm-upstream-status` appears whenever an upstream responded. The other headers are the same as on [POST /v1/chat/completions](/docs/api-chat-completions#response-headers).

## Examples

```bash
curl -N https://api.routerplus.com/v1/messages \
  -H "x-api-key: $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"gpt-4o-mini","stream":true,"max_tokens":100,"messages":[{"role":"user","content":"Say hello."}]}'
```

Yes — a GPT model over the Anthropic wire. Python, with the official `anthropic` SDK and the base URL swapped (the SDK appends `/v1/messages` itself):

```python
import os
from anthropic import Anthropic

client = Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])

with client.messages.stream(
    model="claude-sonnet-4-5",
    max_tokens=300,
    messages=[{"role": "user", "content": "Count to five."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="")
    print("\n", stream.get_final_message().usage)  # tokens, cache splits, cost
```

TypeScript — `@anthropic-ai/sdk`:

```typescript
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  baseURL: "https://api.routerplus.com",
  apiKey: process.env.TM_API_KEY,
});

const stream = client.messages.stream({
  model: "claude-haiku-4-5",
  max_tokens: 200,
  messages: [{ role: "user", content: "Say hello." }],
});
stream.on("text", (t) => process.stdout.write(t));
console.log("\n", (await stream.finalMessage()).usage);
```

Claude Code works unmodified:

```bash
ANTHROPIC_BASE_URL=https://api.routerplus.com ANTHROPIC_AUTH_TOKEN=$TM_API_KEY claude
```

## Model discovery

`GET https://api.routerplus.com/v1/models` returns the catalog; send an `anthropic-version` header (the Anthropic SDK does) to get the Anthropic list shape instead of the OpenAI one. That shape lists chat models only. The public directory with prices and context lengths is `https://app.routerplus.com/api/models.json`.

See also: [POST /v1/chat/completions](/docs/api-chat-completions) · [Errors](/docs/errors) · [Wire compatibility & billing contract](/docs/compat) · [Migrate in one prompt](/docs/migration)


## POST /v1/messages/count_tokens

An unbilled passthrough for Anthropic-dialect deployments — the same request
shape as `/v1/messages`, answered by the provider's own tokenizer:

```bash
curl -s https://api.routerplus.com/v1/messages/count_tokens \
  -H "x-api-key: $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "claude-sonnet-5", "messages": [{"role": "user", "content": "hello"}]}'
```

```json
{ "input_tokens": 8 }
```

The answer is the count only. The call creates no ledger attempt and costs
nothing, but it counts against your key's rate limit and the provider's pool
like any request. An unknown model is a 404; if only OpenAI-dialect deployments
(or a Bedrock connection) serve the model, the response is a typed 400: the
gateway will not fabricate a token estimate.
