Dedicated endpoints
Capacity reserved for your organization on one model: its own limits, contract prices and same-model failover.
A dedicated endpoint is provider capacity reserved for your organization on one model. It has its own limits, prices and failover rules. We set one up with you under a contract; there is no self-serve setup. You call it with the model id we give you, in your organization's own namespace: <your-namespace>/<name>, for example acme/gemma-31b.
Calling an endpoint
Use the endpoint's model id in place of a catalog model id. Everything else stays the same: both surfaces (/v1/chat/completions and /v1/messages), streaming or not, and every request field the model accepts.
curl https://api.routerplus.com/v1/chat/completions \
-H "Authorization: Bearer $TM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "acme/gemma-31b", "messages": [{"role": "user", "content": "Hello"}]}'Only your organization's keys can call the endpoint. It can also be limited to named workspaces. For any other key the id does not exist: it gets the same 404 model_unavailable as an unknown model.
On an open endpoint, the response names the model that ran, for example "model": "google/gemma-4-31b-it", not the endpoint. A closed endpoint names the endpoint (see Closed endpoints). Two headers tell you how the request was served:
| Header | Value |
|---|---|
x-tm-dedicated-endpoint | the endpoint's id, for example acme/gemma-31b |
x-tm-route-role | primary (your dedicated capacity) or fallback (the shared pool) |
In /v1/generation, each attempt's route_context.dedicated records the endpoint id, the revision of its configuration in force, and the role.
GET /v1/models lists your organization's active endpoints, with owned_by routerplus, on the base URL the endpoint is served from. The console's Endpoints page lists them first, before the models you deployed from Optimize: for each endpoint its id, model, base URL, status, provider, price per million tokens (input, cached input and output) and limits.
What serves your request
- 01Your dedicated capacity first.
- 02Then the shared pool, for the same model only. The request moves on when the dedicated deployment fails before your answer starts, when its circuit is open, or when its capacity is full. It goes to public deployments of the same model that meet your endpoint's constraints. It never goes to a different model.
- 03The commit boundary holds. Once the first token reaches you, the request stays with that provider. See Routing & failover.
The constraints are part of your endpoint's configuration and apply to every fallback:
| Constraint | A fallback host qualifies only if |
|---|---|
| precision | it serves one of the listed precisions, for example bf16 |
| zero data retention | it keeps no prompts |
| data collection denied | it neither stores nor trains on requests |
| maximum price | its list price per million tokens is at or below the limit |
| regions | it runs in one of the listed countries |
| hosts | it is one of the named hosts |
| require parameters | it accepts every parameter the request sends |
A host that has not stated a value for a constrained property does not qualify. A fallback that reaches several hosts through one route sends the constraints with the request, so only hosts that meet them serve it. A region constraint excludes such a route, because it makes no promise about where a host runs.
An endpoint can also be set to primary_only. It then never falls back; when the dedicated deployment cannot serve, you get the error.
For one request you can narrow the routes but never widen them. "provider": {"allow_fallbacks": false} keeps the request on your dedicated capacity. On open endpoints, provider.only and provider.ignore filter by provider name, as on any request. Closed endpoints select their providers and hosts through the contract configuration: provider.only, provider.ignore, provider.order, provider.upstream and provider.require_parameters return 400 invalid_request, including empty selectors. Omit those fields on closed endpoints. allow_fallbacks remains supported.
Closed endpoints
An endpoint is open unless your contract says otherwise: like any request, it tells you which deployment and host of the shared pool served you, and your dedicated capacity shows as routerplus. A closed endpoint never names a deployment or a host. Every surface you can read names RouterPlus as the provider, whichever route served the request, and apart from the route role nothing you can read differs between the routes:
| Surface | On a closed endpoint |
|---|---|
x-tm-provider | routerplus |
x-tm-served-by, x-tm-upstream-model, x-tm-upstream-status, x-tm-dropped-params, x-tm-affinity-wait-ms | absent |
x-tm-dedicated-endpoint, x-tm-route-role | as on an open endpoint |
x-tm-remaining-rpm, x-tm-remaining-tpm | counted against the endpoint's own limits only |
| Chat completion body | a fixed set of fields, below; model is the endpoint id |
| Stream chunks | the same fields in each chunk; a provider's SSE comment lines are dropped; a provider's error in the stream ends it with the gateway's own error event, and nothing follows that event |
| Error messages | they name routerplus, never a deployment; a provider's own error body is never relayed |
| Error statuses | a provider's 401, 402, 403 or 404 is a retryable 503 upstream_error; a retry-after is the gateway's own, never a provider's |
| Limits | a refusal because the capacity behind one route is full has x-tm-limit-scope: endpoint and x-tm-limit-id: endpoint:<endpoint id> |
provider.only, provider.ignore, provider.order, provider.upstream | refused with 400 invalid_request |
/v1/generation | deployment is routerplus and model the endpoint id; route_context keeps only dedicated (the endpoint, its revision and the role); admission_context is null, dropped_params is empty, and admission events carry no pool scope id |
/v1/usage | deployment is routerplus and model the endpoint id |
| Console: Overview, Usage, Logs and a request's details | RouterPlus as the provider, no host, the endpoint id as the model, and "contract rate" as the price |
| The playground's receipt | served by RouterPlus |
The records stay masked once an endpoint has been closed: what it served before it closed, and after it is opened again, still reads as above.
The fields of a closed chat completion:
| Level | Fields |
|---|---|
| Top level | id (chatcmpl- and the request id without dashes), object, created (when the gateway received your request), model, choices, usage |
| A choice | index, message (not streamed) or delta (streamed), finish_reason (stop, length, tool_calls, content_filter or function_call; null until a stream ends), logprobs (only when you ask for them) |
message or delta | role, content, tool_calls (id, type, function.name, function.arguments, and index in a stream), function_call, refusal, reasoning_content |
usage | prompt_tokens, completion_tokens, total_tokens, prompt_tokens_details.cached_tokens and completion_tokens_details.reasoning_tokens (both always present, 0 when none), cost |
Every other field is dropped, including any field a provider adds later, and a field without a value is left out (a message keeps content: null beside tool calls). A model's thinking arrives as reasoning_content on every route. A tool call's id is the gateway's own, call_ and 24 hex characters, the same on every chunk of the call and as the tool_use id on /v1/messages. In a stream, the first delta of a choice carries the role, the first delta of a tool call its id, type and name, and the finish and the usage each come in a chunk of their own.
Prices
| Attempt | Price |
|---|---|
| On your dedicated capacity | your contract rate card, per million tokens |
| On a fallback | the model's list price, as for any other request; or your rate card, when your contract sets one price on every route |
A closed endpoint always bills your rate card on every route: a second price would tell you which route served. A fallback is tried only when its list price is known and at or below the endpoint's maximum price, if one is set, whichever price you then pay.
In /v1/generation, an attempt billed at your rate card has inference_price_source customer_rate and a price_snapshot_id of the form dedicated:<endpoint id>:r<revision>, so every charge traces back to the rate card that produced it. usage.cost on the response is the full debit across all attempts, as on every request.
Limits
For the endpoint's traffic, its own limits replace the default organization, workspace, principal and key limits:
| Limit | Counted for |
|---|---|
| requests per minute | the endpoint |
| requests per second (optional): a count per second, or a paced rate | the endpoint |
| tokens per minute | the endpoint |
| concurrent requests | the endpoint |
A limit your organization set itself on /console/limits still applies on top. You can lower what one key or workspace may use; you cannot raise the endpoint's limits.
A refusal is the usual 429 rate_limit (see Rate limits & spend caps), with x-tm-limit-scope: endpoint:
| Case | x-tm-limit-kind | x-tm-limit-id |
|---|---|---|
| Over the endpoint's RPM, RPS, TPM or concurrency | rpm, rps, tpm or concurrency | endpoint:<endpoint id> |
| Above a paced endpoint's hard rate | rps | endpoint:<endpoint id> |
| A paced endpoint's wait would be longer than its bound | rps_queue | endpoint:<endpoint id> |
| The dedicated capacity is full and the endpoint does not spill over to the shared pool | the measure that is full | endpoint:<endpoint id> |
| A closed endpoint: the capacity behind every route it could take is full | the measure that is full | endpoint:<endpoint id> |
| Fallback traffic is over the endpoint's fallback cap | rpm, tpm or concurrency | endpoint-fallback:<endpoint id> |
GET /v1/limits?model=<your-namespace>/<name> returns your key's effective limits on the endpoint and a dedicated object: the endpoint, its revision, the model, the output default, the fallback caps and whether it spills over.
If the endpoint's outage mode is closed, requests are refused while the shared counter store is unreachable. In the default bounded_open mode they are admitted under the endpoint's own limits, divided among the gateways that serve it, for up to 10 minutes; see Limits and capacity.
The paced limit
A limit that counts each calendar second refuses requests in many seconds when traffic arrives at random, even at an average below the limit. A paced limit queues a short burst instead. It has three settings: the paced rate, a hard rate above it, and the longest wait.
- Requests leave admission no faster than the paced rate. A request that must wait waits in the gateway.
x-tm-queue-mson the response says how long, in milliseconds. - A request whose wait would be longer than the bound gets 429,
x-tm-limit-kind: rps_queue. Itsretry-aftersays when its wait would fit. - The hard rate refuses a request at once,
x-tm-limit-kind: rps. It allows a burst as long as the bound (at least one second), so it never refuses a request the queue would hold: with a bound of a second or more, a full queue is what refuses. - The wait comes before the request's own deadline starts, so it does not use the endpoint's total time.
- A waiting request holds no concurrency slot. Its place in the pace stays used even if you disconnect.
For example, with a paced rate of 6 requests a second, a hard rate of 6.5 and a 2 s bound:
| Average rate, arriving at random | Refused (429) | Typical wait |
|---|---|---|
| 4 a second | almost none | about 0.1 s |
| 5 a second | about 0.3 % | about 0.3 s |
| 6 a second | about 4 % | about 0.9 s |
| Steadily above 6 a second | the excess, once the queue is full | up to 2 s |
A queue cannot hold a steady excess: over time the endpoint passes its paced rate.
Timeouts and defaults
| Setting | On other requests | On an endpoint |
|---|---|---|
| Attempts per request | up to 8 | 1 to 8; 3 unless set |
| Total time per request | 450 s | 1 to 600 s |
| Wait for a streamed attempt's response headers before the next route | 20 s | 1 to 120 s |
| Time one attempt may take on a request that is not streamed, before the next route | 450 s | 1 to 600 s, below the total |
max_tokens written into a request that sets none | 4,096 | set per endpoint |
A shorter header wait, or attempt limit for requests that are not streamed, makes failover faster when the dedicated deployment stalls. A stalled request still takes that limit plus the fallback's own time. The attempt limit applies only while another route remains: the last route may run for the rest of the total time.
Edges
POST /v1/routedoes not explain dedicated endpoints yet.- Health circuits are shared by every request to a deployment. Two failures in a row on the dedicated deployment send all its traffic to the fallbacks for 30 seconds. Per-endpoint circuit settings are not available yet.
- We change an endpoint's configuration on request. Each change is a new revision, and it reaches every gateway within seconds.
Markdown source for agents: /docs/dedicated-endpoints.md · index at /llms.txt