# Streaming

Set `"stream": true` and both surfaces stream Server-Sent Events (SSE): the
OpenAI-compatible surface at `POST https://api.routerplus.com/v1/chat/completions` and the
Anthropic-compatible surface at `POST https://api.routerplus.com/v1/messages`.

The images route (`POST https://api.routerplus.com/v1/images/generations`) does not stream:
`stream: true` and `partial_images` are typed 400s there. A video is a job you
poll, not a stream — see [POST /v1/videos](/docs/api-videos).

**Same-dialect routes relay.** When the provider serving your model speaks your
surface's dialect (a Claude model on `/v1/messages`, a GPT model on
`/v1/chat/completions`), the provider's stream is relayed record-for-record with
payloads untouched — provider-specific features (server tool use, web search
results, citations, annotations, `pause_turn`, the matched `stop_sequence`)
reach you exactly as if you called the provider directly. The gateway only
observes (for billing and failover), injects `usage.cost` into the provider's
own usage record, and guarantees a well-formed terminator. One exception: an
explicit `stream_options: {"include_usage": false}` turns the relay off, because
the gateway must then remove the usage chunk it asked the provider for.

**Cross-dialect routes re-encode.** When the provider speaks the other dialect,
the gateway translates its stream into your surface's exact wire framing, and
the shapes below hold.

Before the first byte there is a 20-second headers deadline — a provider that
can't answer in time is failed over invisibly (see
[Routing & failover](/docs/routing)). After that a stream is not cut for being
slow, but every request has a deadline: 450 seconds from dispatch, or the
`timeout_ms` of a [routing policy](/docs/routing-policies) (1–600 s). A stream
still open at the deadline ends with the terminal error event described below.

```bash
curl -N https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5","stream":true,"max_tokens":100,
       "messages":[{"role":"user","content":"Say hello"}]}'
```

Every stream response carries `x-request-id`, `x-tm-provider` (the deployment
that served you), `x-tm-attempts`, and `x-tm-upstream-status` headers, plus
`x-tm-dropped-params`, `x-tm-upstream-model` and `x-tm-served-by` when they
apply (see [Wire compatibility](/docs/compat#response-headers)).

## Frame order — OpenAI surface

Each frame is a complete `data: {json}\n\n` line in the documented
`chat.completion.chunk` shape. The order is fixed: role-priming delta, content
deltas, finish chunk, usage chunk, `data: [DONE]`.

```text
data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}

data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":" there."},"finish_reason":null}]}

data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{},"finish_reason":"stop","native_finish_reason":"end_turn"}]}

data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":9,"total_tokens":21,"prompt_tokens_details":{"cached_tokens":0,"cache_write_tokens":0},"completion_tokens_details":{"reasoning_tokens":0},"cost":0.000114}}

data: [DONE]
```

Notes on this surface:

- The **role-priming delta** (`{"role":"assistant","content":""}`) is emitted as
  soon as the upstream proves alive — it is a liveness signal, not content.
- Tool calls stream as `delta.tool_calls` entries with stable ascending
  `index` values; the arguments arrive as string fragments to concatenate. A
  tool called with no arguments gets one `"{}"` fragment.
- Reasoning models stream their thinking as `delta.reasoning_content`.
- Refusal text streams as `delta.refusal`.
- On a translated stream the finish chunk also carries
  `native_finish_reason`, the provider's own stop reason (here `end_turn`).
- The usage chunk is always sent, zero-filled if the provider reported nothing,
  unless you sent `stream_options.include_usage: false`.
- `[DONE]` is only ever sent after a successful finish — **never after an
  error** (see below).

## Frame order — Anthropic surface

Named SSE events (`event: <name>\ndata: {json}\n\n`) exactly as the Anthropic
Messages API frames them: `message_start`, `content_block_start` /
`content_block_delta` / `content_block_stop`, `message_delta`, `message_stop`.

```text
event: message_start
data: {"type":"message_start","message":{"id":"msg_2f6d8e3b","type":"message","role":"assistant","model":"gpt-4o-mini","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":0,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello there."}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"input_tokens":12,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":9,"cost":0.000007}}

event: message_stop
data: {"type":"message_stop"}
```

Notes on this surface:

- On a translated stream the token counts in `message_start` are zeros; the
  authoritative usage (and `cost`) arrives in the final `message_delta`. Read
  usage there, not from `message_start`. On a relayed stream `message_start`
  carries the provider's own counts.
- Tool calls stream as `tool_use` content blocks with `input_json_delta`
  fragments; thinking streams as `thinking` blocks with `thinking_delta` (and
  `signature_delta` where the provider signs it).
- `input_tokens` excludes cache reads and writes, per the
  [billing contract](/docs/compat) — cache traffic is reported separately in
  `cache_read_input_tokens` and `cache_creation_input_tokens`.

## Usage is always in the stream

You never opt in to stream usage. The gateway injects
`stream_options.include_usage` into every OpenAI-dialect upstream dispatch
itself (without it, OpenAI streams carry no usage at all). An explicit
`stream_options: {"include_usage": false}` is honored: you are billed the same,
but no usage chunk is emitted to you. Every other stream ends
with provider-reported token counts, and every **billed** stream carries
`usage.cost` (USD) — computed with the exact integer micro-USD math the ledger
settles with, covering the full request debit including any attempts that
failed over before your answer started.

> [!TIP]
> `cost` on the wire is the number you are charged. You can recompute it any
> time from the token counts and the public prices at
> [https://app.routerplus.com/api/models.json](https://app.routerplus.com/api/models.json) — settlement
> rounds down, in your favor.

## Keep-alive frames

Slow reasoning models can go quiet for a long time mid-answer, and idle TCP
connections get killed by proxies, load balancers, and some HTTP clients. So
once your stream has started, any silence of 15 seconds or more produces a
keep-alive frame:

- **OpenAI surface:** an SSE comment — `: processing` — which the SSE spec
  requires clients to ignore. SDKs never see it.
- **Anthropic surface:** a native ping event —
  `event: ping` / `data: {"type":"ping"}` — the same frame Anthropic's own API
  sends, which Anthropic SDKs already skip.

Keep-alives are measured from the last real write, so the silence you can
observe is bounded at ~15s, and nothing — keep-alives included — is ever sent
after a terminal event.

## Cancelling a stream

Close the connection (abort the HTTP request) and the gateway immediately
aborts the upstream request, so the provider stops generating.

The billing semantics are deliberate and worth knowing:

- **You are billed for what was generated** up to the abort — the attempt is
  recorded with outcome `cancelled` and the usage streamed so far, visible at
  `GET /v1/generation?id=<x-request-id>`. If the provider reported no usage,
  the output you received is estimated at about four characters per token.
- **A cancellation never counts against the provider's uptime.** You hung up;
  the provider did nothing wrong. Cancelled attempts are excluded from the
  health circuits and the public uptime numbers (see
  [Routing & failover](/docs/routing)).

```typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.routerplus.com/v1",
  apiKey: process.env.TM_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "claude-sonnet-5",
  max_tokens: 4096,
  messages: [{ role: "user", content: "Explain SSE, thoroughly." }],
  stream: true,
});

let chars = 0;
for await (const chunk of stream) {
  chars += chunk.choices[0]?.delta?.content?.length ?? 0;
  if (chars > 2000) {
    stream.controller.abort(); // upstream generation stops; billed to here
    break;
  }
}
```

## When a provider dies mid-stream

If the serving provider fails **after** your answer started, the stream does
not silently end and the request is never restarted on another provider (the
commit boundary — see [Routing & failover](/docs/routing)). Instead you get one
terminal error event inside the already-open 200 stream:

**OpenAI surface** — any known usage (and its `cost`) is emitted first, so the
billed partial attempt stays recomputable from the wire, then a final chunk
with `finish_reason: "error"` and a top-level `error`:

```text
data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":""},"finish_reason":"error"}],"error":{"code":"upstream_error","message":"upstream failure","type":"gateway_error","metadata":{"error_type":"upstream_error"}}}
```

**Anthropic surface** — a native `error` event carrying both the
Anthropic-native type string and the stable canonical `error_type`:

```text
event: error
data: {"type":"error","error":{"type":"api_error","message":"upstream failure","error_type":"upstream_error"}}
```

Nothing follows a terminal event — no `[DONE]`, no further frames. The same
event ends a stream that reaches the request deadline, and a stream whose
provider closed the connection without a proper terminator.

> [!NOTE]
> The Anthropic wire has no legal usage slot outside `message_delta`, so an
> Anthropic-surface stream that dies mid-answer cannot carry its partial usage
> in-band. `GET /v1/generation?id=<x-request-id>` returns the settled numbers
> for exactly this case. This is a documented limitation of the wire format,
> not of the ledger.

The full taxonomy of `error_type` values and what to do about each lives in
[Errors & remediation](/docs/errors).

## SDK streaming examples

Both official SDKs work unmodified — swap the base URL and key.

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])

stream = client.chat.completions.create(
    model="claude-sonnet-5",
    max_tokens=200,
    messages=[{"role": "user", "content": "Write a haiku about failover."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:  # the final usage chunk has an empty choices list
        print(f"\ncost: ${chunk.usage.cost}")
```

```python
import os
import anthropic

client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])

with client.messages.stream(
    model="gpt-4o-mini",  # yes — a GPT model over the Anthropic wire
    max_tokens=200,
    messages=[{"role": "user", "content": "Write a haiku about failover."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()
    print(f"\nstop_reason: {final.stop_reason}")
```
