Streaming
SSE on both surfaces, keep-alives, cancellation semantics.
Set "stream": true and both surfaces stream Server-Sent Events (SSE): the OpenAI-compatible surface at POST https://api.routerplus.com/v1/chat/completions and the Anthropic-compatible surface at POST https://api.routerplus.com/v1/messages.
The images route (POST https://api.routerplus.com/v1/images/generations) does not stream: stream: true and partial_images are typed 400s there. A video is a job you poll, not a stream — see POST /v1/videos.
Same-dialect routes relay. When the provider serving your model speaks your surface's dialect (a Claude model on /v1/messages, a GPT model on /v1/chat/completions), the provider's stream is relayed record-for-record with payloads untouched — provider-specific features (server tool use, web search results, citations, annotations, pause_turn, the matched stop_sequence) reach you exactly as if you called the provider directly. The gateway only observes (for billing and failover), injects usage.cost into the provider's own usage record, and guarantees a well-formed terminator. One exception: an explicit stream_options: {"include_usage": false} turns the relay off, because the gateway must then remove the usage chunk it asked the provider for.
Cross-dialect routes re-encode. When the provider speaks the other dialect, the gateway translates its stream into your surface's exact wire framing, and the shapes below hold.
Before the first byte there is a 20-second headers deadline — a provider that can't answer in time is failed over invisibly (see Routing & failover). After that a stream is not cut for being slow, but every request has a deadline: 450 seconds from dispatch, or the timeout_ms of a routing policy (1–600 s). A stream still open at the deadline ends with the terminal error event described below.
curl -N https://api.routerplus.com/v1/chat/completions \
-H "Authorization: Bearer $TM_API_KEY" \
-H "content-type: application/json" \
-d '{"model":"claude-sonnet-5","stream":true,"max_tokens":100,
"messages":[{"role":"user","content":"Say hello"}]}'Every stream response carries x-request-id, x-tm-provider (the deployment that served you), x-tm-attempts, and x-tm-upstream-status headers, plus x-tm-dropped-params, x-tm-upstream-model and x-tm-served-by when they apply (see Wire compatibility).
Frame order — OpenAI surface
Each frame is a complete data: {json}\n\n line in the documented chat.completion.chunk shape. The order is fixed: role-priming delta, content deltas, finish chunk, usage chunk, data: [DONE].
data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":" there."},"finish_reason":null}]}
data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{},"finish_reason":"stop","native_finish_reason":"end_turn"}]}
data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":9,"total_tokens":21,"prompt_tokens_details":{"cached_tokens":0,"cache_write_tokens":0},"completion_tokens_details":{"reasoning_tokens":0},"cost":0.000114}}
data: [DONE]Notes on this surface:
- The role-priming delta (
{"role":"assistant","content":""}) is emitted as soon as the upstream proves alive — it is a liveness signal, not content. - Tool calls stream as
delta.tool_callsentries with stable ascendingindexvalues; the arguments arrive as string fragments to concatenate. A tool called with no arguments gets one"{}"fragment. - Reasoning models stream their thinking as
delta.reasoning_content. - Refusal text streams as
delta.refusal. - On a translated stream the finish chunk also carries
native_finish_reason, the provider's own stop reason (hereend_turn). - The usage chunk is always sent, zero-filled if the provider reported nothing, unless you sent
stream_options.include_usage: false. [DONE]is only ever sent after a successful finish — never after an error (see below).
Frame order — Anthropic surface
Named SSE events (event: <name>\ndata: {json}\n\n) exactly as the Anthropic Messages API frames them: message_start, content_block_start / content_block_delta / content_block_stop, message_delta, message_stop.
event: message_start
data: {"type":"message_start","message":{"id":"msg_2f6d8e3b","type":"message","role":"assistant","model":"gpt-4o-mini","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":0,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":0}}}
event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello there."}}
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"input_tokens":12,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":9,"cost":0.000007}}
event: message_stop
data: {"type":"message_stop"}Notes on this surface:
- On a translated stream the token counts in
message_startare zeros; the authoritative usage (andcost) arrives in the finalmessage_delta. Read usage there, not frommessage_start. On a relayed streammessage_startcarries the provider's own counts. - Tool calls stream as
tool_usecontent blocks withinput_json_deltafragments; thinking streams asthinkingblocks withthinking_delta(andsignature_deltawhere the provider signs it). input_tokensexcludes cache reads and writes, per the billing contract — cache traffic is reported separately incache_read_input_tokensandcache_creation_input_tokens.
Usage is always in the stream
You never opt in to stream usage. The gateway injects stream_options.include_usage into every OpenAI-dialect upstream dispatch itself (without it, OpenAI streams carry no usage at all). An explicit stream_options: {"include_usage": false} is honored: you are billed the same, but no usage chunk is emitted to you. Every other stream ends with provider-reported token counts, and every billed stream carries usage.cost (USD) — computed with the exact integer micro-USD math the ledger settles with, covering the full request debit including any attempts that failed over before your answer started.
cost on the wire is the number you are charged. You can recompute it any time from the token counts and the public prices at https://app.routerplus.com/api/models.json — settlement rounds down, in your favor.
Keep-alive frames
Slow reasoning models can go quiet for a long time mid-answer, and idle TCP connections get killed by proxies, load balancers, and some HTTP clients. So once your stream has started, any silence of 15 seconds or more produces a keep-alive frame:
- OpenAI surface: an SSE comment —
: processing— which the SSE spec requires clients to ignore. SDKs never see it. - Anthropic surface: a native ping event —
event: ping/data: {"type":"ping"}— the same frame Anthropic's own API sends, which Anthropic SDKs already skip.
Keep-alives are measured from the last real write, so the silence you can observe is bounded at ~15s, and nothing — keep-alives included — is ever sent after a terminal event.
Cancelling a stream
Close the connection (abort the HTTP request) and the gateway immediately aborts the upstream request, so the provider stops generating.
The billing semantics are deliberate and worth knowing:
- You are billed for what was generated up to the abort — the attempt is recorded with outcome
cancelledand the usage streamed so far, visible atGET /v1/generation?id=<x-request-id>. If the provider reported no usage, the output you received is estimated at about four characters per token. - A cancellation never counts against the provider's uptime. You hung up; the provider did nothing wrong. Cancelled attempts are excluded from the health circuits and the public uptime numbers (see Routing & failover).
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.routerplus.com/v1",
apiKey: process.env.TM_API_KEY,
});
const stream = await client.chat.completions.create({
model: "claude-sonnet-5",
max_tokens: 4096,
messages: [{ role: "user", content: "Explain SSE, thoroughly." }],
stream: true,
});
let chars = 0;
for await (const chunk of stream) {
chars += chunk.choices[0]?.delta?.content?.length ?? 0;
if (chars > 2000) {
stream.controller.abort(); // upstream generation stops; billed to here
break;
}
}When a provider dies mid-stream
If the serving provider fails after your answer started, the stream does not silently end and the request is never restarted on another provider (the commit boundary — see Routing & failover). Instead you get one terminal error event inside the already-open 200 stream:
OpenAI surface — any known usage (and its cost) is emitted first, so the billed partial attempt stays recomputable from the wire, then a final chunk with finish_reason: "error" and a top-level error:
data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":""},"finish_reason":"error"}],"error":{"code":"upstream_error","message":"upstream failure","type":"gateway_error","metadata":{"error_type":"upstream_error"}}}Anthropic surface — a native error event carrying both the Anthropic-native type string and the stable canonical error_type:
event: error
data: {"type":"error","error":{"type":"api_error","message":"upstream failure","error_type":"upstream_error"}}Nothing follows a terminal event — no [DONE], no further frames. The same event ends a stream that reaches the request deadline, and a stream whose provider closed the connection without a proper terminator.
The Anthropic wire has no legal usage slot outside message_delta, so an Anthropic-surface stream that dies mid-answer cannot carry its partial usage in-band. GET /v1/generation?id=<x-request-id> returns the settled numbers for exactly this case. This is a documented limitation of the wire format, not of the ledger.
The full taxonomy of error_type values and what to do about each lives in Errors & remediation.
SDK streaming examples
Both official SDKs work unmodified — swap the base URL and key.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])
stream = client.chat.completions.create(
model="claude-sonnet-5",
max_tokens=200,
messages=[{"role": "user", "content": "Write a haiku about failover."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
if chunk.usage: # the final usage chunk has an empty choices list
print(f"\ncost: ${chunk.usage.cost}")import os
import anthropic
client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])
with client.messages.stream(
model="gpt-4o-mini", # yes — a GPT model over the Anthropic wire
max_tokens=200,
messages=[{"role": "user", "content": "Write a haiku about failover."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
final = stream.get_final_message()
print(f"\nstop_reason: {final.stop_reason}")Markdown source for agents: /docs/streaming.md · index at /llms.txt