Console
API reference/POST /v1/messages

POST /v1/messages

The Anthropic-compatible surface.

/llms.txt

The Anthropic-compatible surface. Point the Anthropic SDK (or Claude Code) at https://api.routerplus.com with your marketplace key and it works — including against GPT models. Every chat model in the catalog is callable from this endpoint regardless of which provider serves it; image models answer on POST /v1/images/generations only, and video models on POST /v1/videos only; the gateway translates requests, streams, and errors between dialects. The OpenAI-native twin is POST /v1/chat/completions.

POST https://api.routerplus.com/v1/messages

Query strings are tolerated — clients that append them (Claude Code sends /v1/messages?beta=true) route normally.

Headers

HeaderRequiredBehavior
x-api-keyone of these twoYour marketplace key, tm_vk_... — the Anthropic-native header.
AuthorizationBearer tm_vk_... also accepted, same key.
content-typeyesapplication/json.
anthropic-versionnoAccepted for SDK compatibility; the gateway speaks anthropic-version: 2023-06-01 to Anthropic upstreams regardless of what you send.
anthropic-betanoForwarded when an Anthropic-dialect provider serves the request, for the values that do not change how a token is billed: prompt-caching, token-efficient-tools, fine-grained-tool-streaming, interleaved-thinking, claude-code, oauth and computer-use prefixes. Any other value is dropped and recorded in x-tm-dropped-params as header:anthropic-beta:<value>. No effect on other routes.

A missing or invalid key returns 401 in the Anthropic error shape with error_type: "auth". Optionally send HTTP-Referer and X-Title for content-free app attribution.

Request

JSON body, capped at 10 MB (413 request_too_large above that).

ParameterTypeBehavior
modelstring, requiredA catalog id — any chat model, not just Claude. Unknown ids are an honest 404; an image model is a 400 naming /v1/images/generations, a video model a 400 naming /v1/videos. A deployed endpoint id (tm/<name>-v<n>) is served on POST /v1/chat/completions only and is a 400 here.
max_tokensnumberOutput cap, 1 to 32,768. When you omit it, the gateway sets max_tokens: 4096 on the request it forwards, whichever provider serves it. A value above 32,768 is a 400.
messagesarray, requiredRoles user and assistant. Content: a string, or blocks — text, tool_use (assistant), tool_result (user), thinking / redacted_thinking (assistant). image and document blocks are for a model that takes them (architecture.input_modalities in GET /v1/models): a block the model does not take is a typed 400 naming it, also inside a tool_result. image blocks are a typed 400 when translation to an OpenAI-dialect provider is needed; on Anthropic-dialect routes the messages pass through as sent.
systemstring | arrayA string or text blocks.
streambooleanSSE streaming; see below.
temperature, top_pnumberPassed through.
stop_sequencesstring[]Becomes OpenAI stop on cross-dialect routes.
toolsarray{name, description?, input_schema}. Anthropic server tools (computer use, web search, bash, text editor) pass through on Anthropic-dialect routes and are a typed 400 when an OpenAI-dialect provider would serve the request.
tool_choiceobject{"type":"auto"}, {"type":"any"}, {"type":"none"}, or {"type":"tool","name":"..."}. disable_parallel_tool_use: true becomes parallel_tool_calls: false on OpenAI-dialect routes.
thinking, output_configobjectPass through on Anthropic-dialect routes; a typed 400 when translation to an OpenAI-dialect provider is needed — the gateway does not map them yet and will not drop them silently.
top_k, metadataPass through on Anthropic-dialect routes; dropped and recorded on OpenAI-dialect routes. On a house route, metadata.user_id becomes a hash with your organization; see Sessions and prompt caching.
cache_controlobjectAutomatic prompt caching, on the request or on blocks. Passes to Anthropic and OpenRouter (on Bedrock the request-level field becomes a mark on the last block); dropped and recorded on OpenAI and Azure. A request-level field that Anthropic would refuse is dropped and recorded on every route (see Wire compatibility). On a house route, ttl: "1h" is removed (5 minutes) and recorded.
session_idstringA session id, also accepted as the header x-session-id, x-session-affinity or x-claude-code-session-id. It reaches OpenRouter as session_id and OpenAI or Azure as prompt_cache_key; Anthropic has no session input. See Sessions and prompt caching.
mcp_servers, containerA typed 400 on every route: stripping them would change who runs what.
providerobjectRouting controls read by the gateway and never forwarded: require_parameters, order, only, ignore, allow_fallbacks, upstream and the rest — see Routing policies. An unknown control is a 400.

Any other top-level key is dropped and recorded on every route (the x-tm-dropped-params header names it), never mutated. A top-level api_key, base_url or connection_id is a 400. Unsupported content is never dropped: a block the model does not take, or one a translation cannot carry, is a typed 400 naming the exact field, because dropped content would still be billed upstream. A typed 400 from translation is decided per deployment: when another deployment of the model speaks the Anthropic dialect, the request goes there instead — see Wire compatibility.

Note

Transcripts with thinking blocks round-trip safely. When a model served over the OpenAI dialect streams reasoning, this surface encodes it as thinking blocks — and when you echo that transcript back and the request routes to an OpenAI-dialect provider, assistant thinking/redacted_thinking blocks are dropped rather than rejected. The route that produced your transcript will never 400 on it.

Streamed response

With stream: true, the real Anthropic Messages event sequence — the same normalized encoding whichever provider serves the request. On a translated stream the message id is msg_<request-id>, joinable with the x-request-id header and /v1/generation?id=.

event: message_start
data: {"type":"message_start","message":{"id":"msg_1c9a7b2e-...","type":"message","role":"assistant","model":"claude-sonnet-4-5","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":0,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"input_tokens":12,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":5,"cost":0.000111}}

event: message_stop
data: {"type":"message_stop"}

Block types you may see: text (text_delta), thinking (thinking_delta, plus signature_delta for multi-turn thinking with tools), redacted_thinking, and tool_use (input_json_delta fragments). Streamed stop_reason is one of end_turn, max_tokens, tool_use, refusal.

Usage fields, per the billing contract: input_tokens excludes cache reads and writes; cache_read_input_tokens and cache_creation_input_tokens carry the cache splits; usage.cost (USD, billed keys only) in the final message_delta is the full request debit, failed-over attempts included, computed with the exact math the ledger settles with.

Note

Treat the final message_delta's usage as authoritative. On same-dialect routes (a Claude model on this surface) the stream is relayed record-for-record, so message_start carries the provider's real token counts and its own message id. On cross-dialect routes (an OpenAI-dialect provider behind this surface) the gateway re-encodes the stream and message_start's counts are zeros — the final message_delta always carries the real numbers on every route.

During post-start silences of 15 s or more the gateway emits native ping events ({"type":"ping"}) so idle-but-healthy streams survive proxies. If the serving provider dies after output started, you get one terminal error event inside the 200 stream — carrying the Anthropic-native type plus a stable error_type — and nothing after it; the answer is never silently restarted on another provider. The same event ends a stream that reaches the request deadline (450 s from dispatch, or a routing policy's timeout_ms). One documented limit: a billed stream that dies mid-answer has no legal Anthropic wire slot for usage outside message_delta, so recompute that request's cost from /v1/generation?id=.

Non-streamed response

json
{
  "id": "msg_1c9a7b2e-...",
  "type": "message",
  "role": "assistant",
  "model": "gpt-4o-mini",
  "content": [{ "type": "text", "text": "Hello!" }],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 12,
    "output_tokens": 5,
    "cache_read_input_tokens": 0,
    "cache_creation_input_tokens": 0,
    "cost": 0.000004
  }
}

stop_reason on translated responses is end_turn, max_tokens, tool_use, or refusal; when an Anthropic-dialect provider serves the request the body passes through with its native values (including stop_sequence) plus the injected cost.

Errors

Anthropic error shape with the native type string your SDK switches on (authentication_error, rate_limit_error, billing_error, not_found_error, ...) plus a stable error_type carrying the canonical class:

json
{
  "type": "error",
  "error": {
    "type": "not_found_error",
    "message": "model \"claude-9\" is not in the catalog; GET /v1/models lists what this key can serve",
    "error_type": "model_unavailable"
  },
  "request_id": "..."
}

Out-of-credits and spend-cap breaches are 429 with native type billing_error and error_type: "insufficient_quota" (cap breaches include the exact UTC reset time and an x-tm-cap-reset header). A key whose signup email is not verified yet gets 403 with native type permission_error and error_type: "email_not_verified". Until an organization has added credit, its keys and the organization as a whole run at 20 RPM — 429 rate_limit with retry-after. A provider's own error body is never relayed; when a provider rejected the request, the class says so and x-tm-upstream-status carries its status. Full table: Errors. Every response carries x-request-id; served requests add x-tm-provider and x-tm-attempts, failures add x-tm-error-code (and x-tm-error-origin when a limit, a provider or our infrastructure refused the request), and x-tm-upstream-status appears whenever an upstream responded. The other headers are the same as on POST /v1/chat/completions.

Examples

bash
curl -N https://api.routerplus.com/v1/messages \
  -H "x-api-key: $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"gpt-4o-mini","stream":true,"max_tokens":100,"messages":[{"role":"user","content":"Say hello."}]}'

Yes — a GPT model over the Anthropic wire. Python, with the official anthropic SDK and the base URL swapped (the SDK appends /v1/messages itself):

python
import os
from anthropic import Anthropic

client = Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])

with client.messages.stream(
    model="claude-sonnet-4-5",
    max_tokens=300,
    messages=[{"role": "user", "content": "Count to five."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="")
    print("\n", stream.get_final_message().usage)  # tokens, cache splits, cost
typescript
@anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  baseURL: "https://api.routerplus.com",
  apiKey: process.env.TM_API_KEY,
});

const stream = client.messages.stream({
  model: "claude-haiku-4-5",
  max_tokens: 200,
  messages: [{ role: "user", content: "Say hello." }],
});
stream.on("text", (t) => process.stdout.write(t));
console.log("\n", (await stream.finalMessage()).usage);

Claude Code works unmodified:

bash
ANTHROPIC_BASE_URL=https://api.routerplus.com ANTHROPIC_AUTH_TOKEN=$TM_API_KEY claude

Model discovery

GET https://api.routerplus.com/v1/models returns the catalog; send an anthropic-version header (the Anthropic SDK does) to get the Anthropic list shape instead of the OpenAI one. That shape lists chat models only. The public directory with prices and context lengths is https://app.routerplus.com/api/models.json.

See also: POST /v1/chat/completions · Errors · Wire compatibility & billing contract · Migrate in one prompt

POST /v1/messages/count_tokens

An unbilled passthrough for Anthropic-dialect deployments — the same request shape as /v1/messages, answered by the provider's own tokenizer:

bash
curl -s https://api.routerplus.com/v1/messages/count_tokens \
  -H "x-api-key: $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "claude-sonnet-5", "messages": [{"role": "user", "content": "hello"}]}'
json
{ "input_tokens": 8 }

The answer is the count only. The call creates no ledger attempt and costs nothing, but it counts against your key's rate limit and the provider's pool like any request. An unknown model is a 404; if only OpenAI-dialect deployments (or a Bedrock connection) serve the model, the response is a typed 400: the gateway will not fabricate a token estimate.

Markdown source for agents: /docs/api-messages.md · index at /llms.txt