Console
Core concepts/Streaming

Streaming

SSE on both surfaces, keep-alives, cancellation semantics.

/llms.txt

Set "stream": true and both surfaces stream Server-Sent Events (SSE): the OpenAI-compatible surface at POST https://api.routerplus.com/v1/chat/completions and the Anthropic-compatible surface at POST https://api.routerplus.com/v1/messages.

The images route (POST https://api.routerplus.com/v1/images/generations) does not stream: stream: true and partial_images are typed 400s there. A video is a job you poll, not a stream — see POST /v1/videos.

Same-dialect routes relay. When the provider serving your model speaks your surface's dialect (a Claude model on /v1/messages, a GPT model on /v1/chat/completions), the provider's stream is relayed record-for-record with payloads untouched — provider-specific features (server tool use, web search results, citations, annotations, pause_turn, the matched stop_sequence) reach you exactly as if you called the provider directly. The gateway only observes (for billing and failover), injects usage.cost into the provider's own usage record, and guarantees a well-formed terminator. One exception: an explicit stream_options: {"include_usage": false} turns the relay off, because the gateway must then remove the usage chunk it asked the provider for.

Cross-dialect routes re-encode. When the provider speaks the other dialect, the gateway translates its stream into your surface's exact wire framing, and the shapes below hold.

Before the first byte there is a 20-second headers deadline — a provider that can't answer in time is failed over invisibly (see Routing & failover). After that a stream is not cut for being slow, but every request has a deadline: 450 seconds from dispatch, or the timeout_ms of a routing policy (1–600 s). A stream still open at the deadline ends with the terminal error event described below.

bash
curl -N https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5","stream":true,"max_tokens":100,
       "messages":[{"role":"user","content":"Say hello"}]}'

Every stream response carries x-request-id, x-tm-provider (the deployment that served you), x-tm-attempts, and x-tm-upstream-status headers, plus x-tm-dropped-params, x-tm-upstream-model and x-tm-served-by when they apply (see Wire compatibility).

Frame order — OpenAI surface

Each frame is a complete data: {json}\n\n line in the documented chat.completion.chunk shape. The order is fixed: role-priming delta, content deltas, finish chunk, usage chunk, data: [DONE].

data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}

data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":" there."},"finish_reason":null}]}

data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{},"finish_reason":"stop","native_finish_reason":"end_turn"}]}

data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":9,"total_tokens":21,"prompt_tokens_details":{"cached_tokens":0,"cache_write_tokens":0},"completion_tokens_details":{"reasoning_tokens":0},"cost":0.000114}}

data: [DONE]

Notes on this surface:

  • The role-priming delta ({"role":"assistant","content":""}) is emitted as soon as the upstream proves alive — it is a liveness signal, not content.
  • Tool calls stream as delta.tool_calls entries with stable ascending index values; the arguments arrive as string fragments to concatenate. A tool called with no arguments gets one "{}" fragment.
  • Reasoning models stream their thinking as delta.reasoning_content.
  • Refusal text streams as delta.refusal.
  • On a translated stream the finish chunk also carries native_finish_reason, the provider's own stop reason (here end_turn).
  • The usage chunk is always sent, zero-filled if the provider reported nothing, unless you sent stream_options.include_usage: false.
  • [DONE] is only ever sent after a successful finish — never after an error (see below).

Frame order — Anthropic surface

Named SSE events (event: <name>\ndata: {json}\n\n) exactly as the Anthropic Messages API frames them: message_start, content_block_start / content_block_delta / content_block_stop, message_delta, message_stop.

event: message_start
data: {"type":"message_start","message":{"id":"msg_2f6d8e3b","type":"message","role":"assistant","model":"gpt-4o-mini","content":[],"stop_reason":null,"stop_sequence":null,"usage":{"input_tokens":0,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":0}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Hello there."}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"input_tokens":12,"cache_read_input_tokens":0,"cache_creation_input_tokens":0,"output_tokens":9,"cost":0.000007}}

event: message_stop
data: {"type":"message_stop"}

Notes on this surface:

  • On a translated stream the token counts in message_start are zeros; the authoritative usage (and cost) arrives in the final message_delta. Read usage there, not from message_start. On a relayed stream message_start carries the provider's own counts.
  • Tool calls stream as tool_use content blocks with input_json_delta fragments; thinking streams as thinking blocks with thinking_delta (and signature_delta where the provider signs it).
  • input_tokens excludes cache reads and writes, per the billing contract — cache traffic is reported separately in cache_read_input_tokens and cache_creation_input_tokens.

Usage is always in the stream

You never opt in to stream usage. The gateway injects stream_options.include_usage into every OpenAI-dialect upstream dispatch itself (without it, OpenAI streams carry no usage at all). An explicit stream_options: {"include_usage": false} is honored: you are billed the same, but no usage chunk is emitted to you. Every other stream ends with provider-reported token counts, and every billed stream carries usage.cost (USD) — computed with the exact integer micro-USD math the ledger settles with, covering the full request debit including any attempts that failed over before your answer started.

Tip

cost on the wire is the number you are charged. You can recompute it any time from the token counts and the public prices at https://app.routerplus.com/api/models.json — settlement rounds down, in your favor.

Keep-alive frames

Slow reasoning models can go quiet for a long time mid-answer, and idle TCP connections get killed by proxies, load balancers, and some HTTP clients. So once your stream has started, any silence of 15 seconds or more produces a keep-alive frame:

  • OpenAI surface: an SSE comment — : processing — which the SSE spec requires clients to ignore. SDKs never see it.
  • Anthropic surface: a native ping event — event: ping / data: {"type":"ping"} — the same frame Anthropic's own API sends, which Anthropic SDKs already skip.

Keep-alives are measured from the last real write, so the silence you can observe is bounded at ~15s, and nothing — keep-alives included — is ever sent after a terminal event.

Cancelling a stream

Close the connection (abort the HTTP request) and the gateway immediately aborts the upstream request, so the provider stops generating.

The billing semantics are deliberate and worth knowing:

  • You are billed for what was generated up to the abort — the attempt is recorded with outcome cancelled and the usage streamed so far, visible at GET /v1/generation?id=<x-request-id>. If the provider reported no usage, the output you received is estimated at about four characters per token.
  • A cancellation never counts against the provider's uptime. You hung up; the provider did nothing wrong. Cancelled attempts are excluded from the health circuits and the public uptime numbers (see Routing & failover).
typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.routerplus.com/v1",
  apiKey: process.env.TM_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "claude-sonnet-5",
  max_tokens: 4096,
  messages: [{ role: "user", content: "Explain SSE, thoroughly." }],
  stream: true,
});

let chars = 0;
for await (const chunk of stream) {
  chars += chunk.choices[0]?.delta?.content?.length ?? 0;
  if (chars > 2000) {
    stream.controller.abort(); // upstream generation stops; billed to here
    break;
  }
}

When a provider dies mid-stream

If the serving provider fails after your answer started, the stream does not silently end and the request is never restarted on another provider (the commit boundary — see Routing & failover). Instead you get one terminal error event inside the already-open 200 stream:

OpenAI surface — any known usage (and its cost) is emitted first, so the billed partial attempt stays recomputable from the wire, then a final chunk with finish_reason: "error" and a top-level error:

data: {"id":"chatcmpl-2f6d8e3b","object":"chat.completion.chunk","created":1757000000,"model":"claude-sonnet-5","choices":[{"index":0,"delta":{"content":""},"finish_reason":"error"}],"error":{"code":"upstream_error","message":"upstream failure","type":"gateway_error","metadata":{"error_type":"upstream_error"}}}

Anthropic surface — a native error event carrying both the Anthropic-native type string and the stable canonical error_type:

event: error
data: {"type":"error","error":{"type":"api_error","message":"upstream failure","error_type":"upstream_error"}}

Nothing follows a terminal event — no [DONE], no further frames. The same event ends a stream that reaches the request deadline, and a stream whose provider closed the connection without a proper terminator.

Note

The Anthropic wire has no legal usage slot outside message_delta, so an Anthropic-surface stream that dies mid-answer cannot carry its partial usage in-band. GET /v1/generation?id=<x-request-id> returns the settled numbers for exactly this case. This is a documented limitation of the wire format, not of the ledger.

The full taxonomy of error_type values and what to do about each lives in Errors & remediation.

SDK streaming examples

Both official SDKs work unmodified — swap the base URL and key.

python
import os
from openai import OpenAI

client = OpenAI(base_url="https://api.routerplus.com/v1", api_key=os.environ["TM_API_KEY"])

stream = client.chat.completions.create(
    model="claude-sonnet-5",
    max_tokens=200,
    messages=[{"role": "user", "content": "Write a haiku about failover."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:  # the final usage chunk has an empty choices list
        print(f"\ncost: ${chunk.usage.cost}")
python
import os
import anthropic

client = anthropic.Anthropic(base_url="https://api.routerplus.com", api_key=os.environ["TM_API_KEY"])

with client.messages.stream(
    model="gpt-4o-mini",  # yes — a GPT model over the Anthropic wire
    max_tokens=200,
    messages=[{"role": "user", "content": "Write a haiku about failover."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()
    print(f"\nstop_reason: {final.stop_reason}")

Markdown source for agents: /docs/streaming.md · index at /llms.txt