# POST /v1/chat/completions

The OpenAI-compatible surface. Point any OpenAI SDK (or OpenAI-compatible client) at `https://api.routerplus.com/v1` with your marketplace key and it works — including against Claude models. Every chat model in the catalog is callable from this endpoint no matter which provider serves it; image models answer on [POST /v1/images/generations](/docs/api-images) only, and video models on [POST /v1/videos](/docs/api-videos) only; the gateway translates requests, streams, and errors between dialects. The Anthropic-native twin of this endpoint is [POST /v1/messages](/docs/api-messages).

```
POST https://api.routerplus.com/v1/chat/completions
```

## Authentication

| Header | Format |
|---|---|
| `Authorization` | `Bearer tm_vk_...` (the standard OpenAI-style header) |
| `x-api-key` | `tm_vk_...` — also accepted, same key |

A missing or invalid key returns 401 with `error_type: "auth"`. Optionally send `HTTP-Referer` and `X-Title` to identify your app (content-free attribution).

## Request

JSON body, capped at 10 MB (413 `request_too_large` above that).

| Parameter | Type | Behavior |
|---|---|---|
| `model` | string, required | A catalog id — see [Model discovery](#model-discovery). Unknown ids are an honest 404, never a silent substitute. A deployed endpoint from [Optimize](/docs/model-search) (`tm/<name>-v<n>`) is served here too, with a key of the organization that owns it, and so is a [dedicated endpoint](/docs/dedicated-endpoints) (`<your-namespace>/<name>`). |
| `messages` | array, required | Roles `system`, `developer`, `user`, `assistant`, `tool`. Content: a string, or `text` parts, plus `image_url`, `file`, `input_audio` and `video_url` parts for a model that takes them (`architecture.input_modalities` in [`GET /v1/models`](/docs/api-models)). A part the model does not take is a typed 400 naming the part. Media parts reach the provider only when an OpenAI-dialect provider serves the request; on a route that needs translation they are a typed 400 naming the part. |
| `stream` | boolean | Server-sent events; see [Streamed response](#streamed-response). |
| `max_tokens` | number | Output cap, 1 to 32,768. **When you send neither `max_tokens` nor `max_completion_tokens`, the gateway sets `max_tokens: 4096`** on the request it forwards, whichever provider serves it. A value above 32,768 is a 400. |
| `max_completion_tokens` | number | Honored as an alias. Sending both with different values is a 400. |
| `temperature` | number | Passed through. Above 1 it cannot be translated to an Anthropic-dialect provider — see [Wire compatibility](/docs/compat). |
| `top_p` | number | Passed through. |
| `stop` | string \| string[] | Becomes Anthropic `stop_sequences` on cross-dialect routes. |
| `tools` | array | `type: "function"` tools only. A tool without `parameters` gets an empty object schema on Anthropic-dialect routes. |
| `tool_choice` | string \| object | `"auto"`, `"none"`, `"required"`, or `{"type":"function","function":{"name":"..."}}`. |
| `parallel_tool_calls` | boolean | `false` becomes `disable_parallel_tool_use` on Anthropic-dialect routes. |
| `n` | number | **Rejected when > 1** — 400 `invalid_request`. See below. |
| `stream_options` | object | Usage is always on. An explicit `{"include_usage": false}` is honored: you are billed the same, but no usage chunk is written to you. Otherwise the object reaches an OpenAI-dialect provider with `include_usage` set, and is dropped and recorded on a translated route. |
| `provider` | object | Routing controls read by the gateway and never forwarded: `require_parameters`, `order`, `only`, `ignore`, `allow_fallbacks`, `upstream` and the rest — see [Routing policies](/docs/routing-policies). An unknown control is a 400. |
| `session_id` | string | A session id: the gateway sends each provider its own session field (`prompt_cache_key` to OpenAI and Azure, `session_id` to OpenRouter). The headers `x-session-id`, `x-session-affinity` and `x-claude-code-session-id` work too. See [Sessions and prompt caching](/docs/compat#sessions-and-prompt-caching). |
| `prompt_cache_key` | string | OpenAI's cache key: sent unchanged to OpenAI and Azure, and to OpenRouter. |
| `prompt_cache_retention` | string | `"in_memory"` or `"24h"`: sent unchanged to OpenAI and Azure; dropped and recorded elsewhere. |
| `safety_identifier` | string | Your end-user id. On a house route the gateway sends a hash of it with your organization; see [Sessions and prompt caching](/docs/compat#sessions-and-prompt-caching). |
| `cache_control` | object | Automatic prompt caching, as on the Anthropic API: sent to Anthropic and OpenRouter; dropped and recorded on OpenAI and Azure. A field that Anthropic would refuse is dropped and recorded on every route (see [Wire compatibility](/docs/compat)). A `cache_control` mark on a `text` part is kept for Anthropic and OpenRouter, and passes unchanged to OpenAI and Azure, which ignore it. |

A top-level `api_key`, `base_url`, `connection_id` or similar credential or destination field is a 400: keys and endpoints are configured in the console, never sent in a request.

### Everything else

The table above is the **portable** set — it survives translation to any provider. What happens to parameters outside it depends on who serves the request: when the serving provider speaks the OpenAI dialect, the OpenAI chat parameters (`response_format`, `reasoning_effort`, `seed`, `user`, `logprobs`, `top_logprobs`, `frequency_penalty`, `presence_penalty`, `logit_bias`, `metadata`, `store`, `prediction`) reach it as you sent them; when translation to the Anthropic dialect is needed, `response_format`, `reasoning_effort`, `prediction`, `audio` and `modalities` are a typed 400 (they change what the model produces), and other unknown parameters are dropped — never silently mutated into something else. Any top-level key outside those sets is dropped and recorded on every route. Every drop is named in the `x-tm-dropped-params` header; `"provider": {"require_parameters": true}` turns a drop into a 400 instead. The full tables are on [Wire compatibility](/docs/compat).

Unsupported **content** is different from unsupported parameters: it is never dropped, because dropped content would still be billed upstream. A part the model does not take — an image to a model that takes text only, say — is a typed 400 naming the exact field (e.g. `messages[0].content[1]`), on every route, before anything is billed or sent. On routes that need translation to the Anthropic dialect, image parts, audio parts, and unknown content-part types are a typed 400 naming the field too. When the serving provider speaks the OpenAI dialect and the model takes the part, content passes through as sent.

### Why `n > 1` is rejected

The gateway's normalizer emits exactly one choice. Accepting `n: 3` and billing you for a garbled single-choice response would be dishonest, so the request is refused up front:

```json
{
  "error": {
    "code": "invalid_request",
    "message": "n>1 is not supported in v1; request a single choice",
    "type": "invalid_request_error",
    "metadata": { "error_type": "invalid_request" }
  },
  "request_id": "..."
}
```

## Non-streamed response

A standard `chat.completion` object. On translated responses the `id` is `chatcmpl-<request-id>`; when an OpenAI-dialect provider serves the request, the body passes through with the provider's own `id`. For the `/v1/generation?id=` audit, use the `x-request-id` response header — it carries the bare request id on every route.

On a [closed dedicated endpoint](/docs/dedicated-endpoints#closed-endpoints) the body keeps a fixed set of fields on every route: `model` is the endpoint id, the `id` is `chatcmpl-` and the request id without dashes, and provider fields such as `system_fingerprint` and `native_finish_reason` are dropped.

```json
{
  "id": "chatcmpl-1c9a7b2e-...",
  "object": "chat.completion",
  "created": 1757000000,
  "model": "claude-sonnet-4-5",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Hello!" },
      "finish_reason": "stop",
      "native_finish_reason": "end_turn"
    }
  ],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 5,
    "total_tokens": 17,
    "prompt_tokens_details": { "cached_tokens": 0, "cache_write_tokens": 0 },
    "cost": 0.000111
  }
}
```

Usage fields (the billing contract — see [Wire compatibility](/docs/compat)):

| Field | Meaning |
|---|---|
| `prompt_tokens` | All billed input tokens — **cache-inclusive**: uncached input + cache reads + cache writes. |
| `prompt_tokens_details.cached_tokens` | Cache reads (billed at the cached rate). |
| `prompt_tokens_details.cache_write_tokens` | Cache writes (billed at their own rate). Always present on billed responses, 0 when none. |
| `completion_tokens` | Output tokens. Reasoning tokens are a subset, reported in `completion_tokens_details.reasoning_tokens` when the provider reports them (streams always carry the field). |
| `cost` | Settled cost in USD — billed keys only. Computed with the exact integer micro-USD math the ledger settles with, covering the **full request** debit including attempts that failed over before your answer started. You can recompute your bill from the wire. |

`finish_reason` is one of `stop`, `length`, `tool_calls`, `content_filter`. On responses translated from an Anthropic-dialect provider, `choices[0].native_finish_reason` preserves the provider-raw stop reason (e.g. `end_turn`, `stop_sequence`) — a documented extension. A model's thinking, when the provider returns it, is in `message.reasoning_content`.

## Streamed response

With `stream: true`, SSE frames in `chat.completion.chunk` shape, in a fixed order:

1. A role-priming delta (`{"role":"assistant","content":""}`) as soon as the upstream proves alive.
2. Content deltas: `delta.content`, `delta.reasoning_content` (thinking models), and `delta.tool_calls` fragments addressed by `index` (id and `function.name` arrive on each call's first fragment).
3. A finish chunk with `finish_reason` set, plus `native_finish_reason` on translated streams.
4. A usage chunk (empty `choices`) — always sent, no `stream_options` needed:

```text
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":1757000000,"model":"claude-haiku-4-5","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":5,"total_tokens":17,"prompt_tokens_details":{"cached_tokens":0,"cache_write_tokens":0},"completion_tokens_details":{"reasoning_tokens":0},"cost":0.000037}}

data: [DONE]
```

5. `data: [DONE]`.

When an OpenAI-dialect provider serves the request, its stream is relayed record for record, with `usage.cost` added to its usage chunk; the shapes above are what a translated stream looks like. On a closed dedicated endpoint each chunk keeps only the fixed set of fields, and a provider's own SSE comment lines are dropped. Full detail on [Streaming](/docs/streaming).

During post-start silences of 15 s or more (slow reasoning models), the gateway emits an SSE comment (`: processing`) so proxies don't kill an idle-but-healthy stream — spec-ignorable, and your SDK already ignores it.

> [!WARNING]
> If the serving provider dies **after** your answer started, the stream carries one terminal error chunk — a full chunk envelope with `finish_reason: "error"` and a top-level `error` object carrying `metadata.error_type` — and nothing after it. **No `[DONE]` follows an error.** The answer is never silently restarted on a different provider mid-stream; failover happens only before the first byte of output, invisibly. The same chunk ends a stream that reaches the request deadline (450 s from dispatch, or a routing policy's `timeout_ms`).

## Errors

Every error uses the OpenAI envelope with a stable `error.metadata.error_type` (`auth`, `rate_limit`, `insufficient_quota`, `model_unavailable`, `invalid_request`, `context_overflow`, `content_policy`, `upstream_error`, `gateway_error`, ...). A provider's own error body is never relayed; when a provider rejected the request, the class says so and the `x-tm-upstream-status` header carries its status. Full table with remediations: [Errors](/docs/errors).

Out-of-credits is the OpenAI-standard **429 `insufficient_quota`** (your SDK already recognizes it), covering an empty balance or a monthly spend cap — cap breaches include the exact UTC reset time in the message and the `x-tm-cap-reset` header. First trial keys shown after verified sign-in have a 20 RPM limit, and until an organization buys credit its keys together run at 20 RPM; 429 `rate_limit` carries `retry-after`.

> [!TIP]
> Billed requests reserve worst-case cost before dispatch: the request's size in bytes as the input estimate (one token per byte, plus 65,536 per image part; an image sent inline as base64 counts the 65,536 only, not its bytes) and `max_tokens` (4096 when omitted) as the output, at the model's rates, rounded up. If your balance can't cover it — including your org's other in-flight reservations — you get `insufficient_quota` before any provider is called. Settlement is on observed usage only, rounded down, so a generous `max_tokens` never costs extra — but it can make a low balance refuse early. Set it realistically.

## Response headers

| Header | Meaning |
|---|---|
| `x-request-id` | Gateway request id; joins `/v1/usage` and `/v1/generation?id=`. |
| `x-tm-provider` | Deployment that served the request; `routerplus` on a closed dedicated endpoint. |
| `x-tm-attempts` | Physical dispatches, failovers included. |
| `x-tm-upstream-status` | The provider's own HTTP status. |
| `x-tm-error-code` | Canonical error class, on failures. `x-tm-error-origin` says where the failure came from, and `x-tm-limit-scope` / `x-tm-limit-kind` name a limit that refused the request — see [Errors](/docs/errors). |
| `x-tm-upstream-model` | Present when the serving provider knows this model under its own id (aggregators): the id the gateway actually sent. The response's `model` field echoes the provider's id. Absent on a closed dedicated endpoint. |
| `x-tm-served-by` | Present when the serving route is an aggregator that names the provider it used (for example `Amazon Bedrock`). Absent on a direct route: there `x-tm-provider` is the provider. Absent on a closed dedicated endpoint. |
| `x-tm-dedicated-endpoint`, `x-tm-route-role` | On a [dedicated endpoint](/docs/dedicated-endpoints): its id, and `primary` or `fallback`. |
| `x-tm-queue-ms` | On a dedicated endpoint with a [paced limit](/docs/dedicated-endpoints#the-paced-limit): how long the request waited for its turn, in milliseconds. |
| `x-tm-dropped-params` | Comma-joined field paths of any parameters the gateway stripped for this route (D8 §2: drops are recorded, never silent). Absent when nothing was dropped. Send `"provider": {"require_parameters": true}` to get a typed 400 instead of any drop. |
| `x-tm-route-plan-id` | On billed requests: the id of the route plan that chose the deployments. |
| `x-tm-admission-mode`, `x-tm-remaining-rpm`, `x-tm-remaining-tpm` | On billed requests: how the request was admitted and the smallest allowance left across the limits it claimed — see [Limits and capacity](/docs/admission). |

## Examples

```bash
curl -N https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"claude-haiku-4-5","stream":true,"max_tokens":100,"messages":[{"role":"user","content":"Say hello."}]}'
```

Python — the official `openai` SDK with the base URL swapped:

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])

resp = client.chat.completions.create(
    model="claude-sonnet-4-5",  # a Claude model over the OpenAI wire — the gateway translates
    max_tokens=200,
    messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(resp.choices[0].message.content)
print(resp.usage)  # prompt/completion tokens, cache splits, cost (USD)
```

TypeScript — the official `openai` package:

```typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.routerplus.com/v1",
  apiKey: process.env.TM_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "gpt-4o-mini",
  stream: true,
  messages: [{ role: "user", content: "Count to five." }],
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
  if (chunk.usage) console.log("\n", chunk.usage); // final chunk: tokens, cache splits, cost
}
```

## Model discovery

`GET https://api.routerplus.com/v1/models` (authenticated) returns the catalog in the OpenAI list shape — one row per model id — and your organization's dedicated endpoints. The public directory with prices and context lengths is `https://app.routerplus.com/api/models.json`. A model id that is not listed returns 404 `model_unavailable` naming the id you asked for; if every deployment serving a listed model is cooling down after failures, you get a 503 `gateway_error` with `retry-after: 5` and `x-tm-limit-kind: health` instead of a fake 404 — retry, don't fix.

See also: [POST /v1/messages](/docs/api-messages) · [Errors](/docs/errors) · [Wire compatibility & billing contract](/docs/compat) · [Quickstart](/docs/quickstart)
