Console
Get started/Dedicated endpoints

Dedicated endpoints

Capacity reserved for your organization on one model: its own limits, contract prices and same-model failover.

/llms.txt

A dedicated endpoint is provider capacity reserved for your organization on one model. It has its own limits, prices and failover rules. We set one up with you under a contract; there is no self-serve setup. You call it with the model id we give you, in your organization's own namespace: <your-namespace>/<name>, for example acme/gemma-31b.

Calling an endpoint

Use the endpoint's model id in place of a catalog model id. Everything else stays the same: both surfaces (/v1/chat/completions and /v1/messages), streaming or not, and every request field the model accepts.

bash
curl https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "acme/gemma-31b", "messages": [{"role": "user", "content": "Hello"}]}'

Only your organization's keys can call the endpoint. It can also be limited to named workspaces. For any other key the id does not exist: it gets the same 404 model_unavailable as an unknown model.

On an open endpoint, the response names the model that ran, for example "model": "google/gemma-4-31b-it", not the endpoint. A closed endpoint names the endpoint (see Closed endpoints). Two headers tell you how the request was served:

HeaderValue
x-tm-dedicated-endpointthe endpoint's id, for example acme/gemma-31b
x-tm-route-roleprimary (your dedicated capacity) or fallback (the shared pool)

In /v1/generation, each attempt's route_context.dedicated records the endpoint id, the revision of its configuration in force, and the role.

GET /v1/models lists your organization's active endpoints, with owned_by routerplus, on the base URL the endpoint is served from. The console's Endpoints page lists them first, before the models you deployed from Optimize: for each endpoint its id, model, base URL, status, provider, price per million tokens (input, cached input and output) and limits.

What serves your request

  1. 01
    Your dedicated capacity first.
  2. 02
    Then the shared pool, for the same model only. The request moves on when the dedicated deployment fails before your answer starts, when its circuit is open, or when its capacity is full. It goes to public deployments of the same model that meet your endpoint's constraints. It never goes to a different model.
  3. 03
    The commit boundary holds. Once the first token reaches you, the request stays with that provider. See Routing & failover.

The constraints are part of your endpoint's configuration and apply to every fallback:

ConstraintA fallback host qualifies only if
precisionit serves one of the listed precisions, for example bf16
zero data retentionit keeps no prompts
data collection deniedit neither stores nor trains on requests
maximum priceits list price per million tokens is at or below the limit
regionsit runs in one of the listed countries
hostsit is one of the named hosts
require parametersit accepts every parameter the request sends

A host that has not stated a value for a constrained property does not qualify. A fallback that reaches several hosts through one route sends the constraints with the request, so only hosts that meet them serve it. A region constraint excludes such a route, because it makes no promise about where a host runs.

An endpoint can also be set to primary_only. It then never falls back; when the dedicated deployment cannot serve, you get the error.

For one request you can narrow the routes but never widen them. "provider": {"allow_fallbacks": false} keeps the request on your dedicated capacity. On open endpoints, provider.only and provider.ignore filter by provider name, as on any request. Closed endpoints select their providers and hosts through the contract configuration: provider.only, provider.ignore, provider.order, provider.upstream and provider.require_parameters return 400 invalid_request, including empty selectors. Omit those fields on closed endpoints. allow_fallbacks remains supported.

Closed endpoints

An endpoint is open unless your contract says otherwise: like any request, it tells you which deployment and host of the shared pool served you, and your dedicated capacity shows as routerplus. A closed endpoint never names a deployment or a host. Every surface you can read names RouterPlus as the provider, whichever route served the request, and apart from the route role nothing you can read differs between the routes:

SurfaceOn a closed endpoint
x-tm-providerrouterplus
x-tm-served-by, x-tm-upstream-model, x-tm-upstream-status, x-tm-dropped-params, x-tm-affinity-wait-msabsent
x-tm-dedicated-endpoint, x-tm-route-roleas on an open endpoint
x-tm-remaining-rpm, x-tm-remaining-tpmcounted against the endpoint's own limits only
Chat completion bodya fixed set of fields, below; model is the endpoint id
Stream chunksthe same fields in each chunk; a provider's SSE comment lines are dropped; a provider's error in the stream ends it with the gateway's own error event, and nothing follows that event
Error messagesthey name routerplus, never a deployment; a provider's own error body is never relayed
Error statusesa provider's 401, 402, 403 or 404 is a retryable 503 upstream_error; a retry-after is the gateway's own, never a provider's
Limitsa refusal because the capacity behind one route is full has x-tm-limit-scope: endpoint and x-tm-limit-id: endpoint:<endpoint id>
provider.only, provider.ignore, provider.order, provider.upstreamrefused with 400 invalid_request
/v1/generationdeployment is routerplus and model the endpoint id; route_context keeps only dedicated (the endpoint, its revision and the role); admission_context is null, dropped_params is empty, and admission events carry no pool scope id
/v1/usagedeployment is routerplus and model the endpoint id
Console: Overview, Usage, Logs and a request's detailsRouterPlus as the provider, no host, the endpoint id as the model, and "contract rate" as the price
The playground's receiptserved by RouterPlus

The records stay masked once an endpoint has been closed: what it served before it closed, and after it is opened again, still reads as above.

The fields of a closed chat completion:

LevelFields
Top levelid (chatcmpl- and the request id without dashes), object, created (when the gateway received your request), model, choices, usage
A choiceindex, message (not streamed) or delta (streamed), finish_reason (stop, length, tool_calls, content_filter or function_call; null until a stream ends), logprobs (only when you ask for them)
message or deltarole, content, tool_calls (id, type, function.name, function.arguments, and index in a stream), function_call, refusal, reasoning_content
usageprompt_tokens, completion_tokens, total_tokens, prompt_tokens_details.cached_tokens and completion_tokens_details.reasoning_tokens (both always present, 0 when none), cost

Every other field is dropped, including any field a provider adds later, and a field without a value is left out (a message keeps content: null beside tool calls). A model's thinking arrives as reasoning_content on every route. A tool call's id is the gateway's own, call_ and 24 hex characters, the same on every chunk of the call and as the tool_use id on /v1/messages. In a stream, the first delta of a choice carries the role, the first delta of a tool call its id, type and name, and the finish and the usage each come in a chunk of their own.

Prices

AttemptPrice
On your dedicated capacityyour contract rate card, per million tokens
On a fallbackthe model's list price, as for any other request; or your rate card, when your contract sets one price on every route

A closed endpoint always bills your rate card on every route: a second price would tell you which route served. A fallback is tried only when its list price is known and at or below the endpoint's maximum price, if one is set, whichever price you then pay.

In /v1/generation, an attempt billed at your rate card has inference_price_source customer_rate and a price_snapshot_id of the form dedicated:<endpoint id>:r<revision>, so every charge traces back to the rate card that produced it. usage.cost on the response is the full debit across all attempts, as on every request.

Limits

For the endpoint's traffic, its own limits replace the default organization, workspace, principal and key limits:

LimitCounted for
requests per minutethe endpoint
requests per second (optional): a count per second, or a paced ratethe endpoint
tokens per minutethe endpoint
concurrent requeststhe endpoint

A limit your organization set itself on /console/limits still applies on top. You can lower what one key or workspace may use; you cannot raise the endpoint's limits.

A refusal is the usual 429 rate_limit (see Rate limits & spend caps), with x-tm-limit-scope: endpoint:

Casex-tm-limit-kindx-tm-limit-id
Over the endpoint's RPM, RPS, TPM or concurrencyrpm, rps, tpm or concurrencyendpoint:<endpoint id>
Above a paced endpoint's hard raterpsendpoint:<endpoint id>
A paced endpoint's wait would be longer than its boundrps_queueendpoint:<endpoint id>
The dedicated capacity is full and the endpoint does not spill over to the shared poolthe measure that is fullendpoint:<endpoint id>
A closed endpoint: the capacity behind every route it could take is fullthe measure that is fullendpoint:<endpoint id>
Fallback traffic is over the endpoint's fallback caprpm, tpm or concurrencyendpoint-fallback:<endpoint id>

GET /v1/limits?model=<your-namespace>/<name> returns your key's effective limits on the endpoint and a dedicated object: the endpoint, its revision, the model, the output default, the fallback caps and whether it spills over.

If the endpoint's outage mode is closed, requests are refused while the shared counter store is unreachable. In the default bounded_open mode they are admitted under the endpoint's own limits, divided among the gateways that serve it, for up to 10 minutes; see Limits and capacity.

The paced limit

A limit that counts each calendar second refuses requests in many seconds when traffic arrives at random, even at an average below the limit. A paced limit queues a short burst instead. It has three settings: the paced rate, a hard rate above it, and the longest wait.

  • Requests leave admission no faster than the paced rate. A request that must wait waits in the gateway. x-tm-queue-ms on the response says how long, in milliseconds.
  • A request whose wait would be longer than the bound gets 429, x-tm-limit-kind: rps_queue. Its retry-after says when its wait would fit.
  • The hard rate refuses a request at once, x-tm-limit-kind: rps. It allows a burst as long as the bound (at least one second), so it never refuses a request the queue would hold: with a bound of a second or more, a full queue is what refuses.
  • The wait comes before the request's own deadline starts, so it does not use the endpoint's total time.
  • A waiting request holds no concurrency slot. Its place in the pace stays used even if you disconnect.

For example, with a paced rate of 6 requests a second, a hard rate of 6.5 and a 2 s bound:

Average rate, arriving at randomRefused (429)Typical wait
4 a secondalmost noneabout 0.1 s
5 a secondabout 0.3 %about 0.3 s
6 a secondabout 4 %about 0.9 s
Steadily above 6 a secondthe excess, once the queue is fullup to 2 s

A queue cannot hold a steady excess: over time the endpoint passes its paced rate.

Timeouts and defaults

SettingOn other requestsOn an endpoint
Attempts per requestup to 81 to 8; 3 unless set
Total time per request450 s1 to 600 s
Wait for a streamed attempt's response headers before the next route20 s1 to 120 s
Time one attempt may take on a request that is not streamed, before the next route450 s1 to 600 s, below the total
max_tokens written into a request that sets none4,096set per endpoint

A shorter header wait, or attempt limit for requests that are not streamed, makes failover faster when the dedicated deployment stalls. A stalled request still takes that limit plus the fallback's own time. The attempt limit applies only while another route remains: the last route may run for the rest of the total time.

Edges

  • POST /v1/route does not explain dedicated endpoints yet.
  • Health circuits are shared by every request to a deployment. Two failures in a row on the dedicated deployment send all its traffic to the fallbacks for 30 seconds. Per-endpoint circuit settings are not available yet.
  • We change an endpoint's configuration on request. Each change is a new revision, and it reaches every gateway within seconds.

Markdown source for agents: /docs/dedicated-endpoints.md · index at /llms.txt