# Dedicated endpoints

A dedicated endpoint is provider capacity reserved for your organization on one model. It has its own limits, prices and failover rules. We set one up with you under a contract; there is no self-serve setup. You call it with the model id we give you, in your organization's own namespace: `<your-namespace>/<name>`, for example `acme/gemma-31b`.

## Calling an endpoint

Use the endpoint's model id in place of a catalog model id. Everything else stays the same: both surfaces (`/v1/chat/completions` and `/v1/messages`), streaming or not, and every request field the model accepts.

```bash
curl https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "acme/gemma-31b", "messages": [{"role": "user", "content": "Hello"}]}'
```

Only your organization's keys can call the endpoint. It can also be limited to named workspaces. For any other key the id does not exist: it gets the same 404 `model_unavailable` as an unknown model.

On an open endpoint, the response names the model that ran, for example `"model": "google/gemma-4-31b-it"`, not the endpoint. A closed endpoint names the endpoint (see [Closed endpoints](#closed-endpoints)). Two headers tell you how the request was served:

| Header | Value |
|---|---|
| `x-tm-dedicated-endpoint` | the endpoint's id, for example `acme/gemma-31b` |
| `x-tm-route-role` | `primary` (your dedicated capacity) or `fallback` (the shared pool) |

In [`/v1/generation`](/docs/api-usage), each attempt's `route_context.dedicated` records the endpoint id, the revision of its configuration in force, and the role.

`GET /v1/models` lists your organization's active endpoints, with `owned_by` `routerplus`, on the base URL the endpoint is served from. The console's [Endpoints](https://app.routerplus.com/endpoints) page lists them first, before the models you deployed from Optimize: for each endpoint its id, model, base URL, status, provider, price per million tokens (input, cached input and output) and limits.

## What serves your request

1. **Your dedicated capacity first.**
2. **Then the shared pool, for the same model only.** The request moves on when the dedicated deployment fails before your answer starts, when its circuit is open, or when its capacity is full. It goes to public deployments of the same model that meet your endpoint's constraints. It never goes to a different model.
3. **The commit boundary holds.** Once the first token reaches you, the request stays with that provider. See [Routing & failover](/docs/routing).

The constraints are part of your endpoint's configuration and apply to every fallback:

| Constraint | A fallback host qualifies only if |
|---|---|
| precision | it serves one of the listed precisions, for example `bf16` |
| zero data retention | it keeps no prompts |
| data collection denied | it neither stores nor trains on requests |
| maximum price | its list price per million tokens is at or below the limit |
| regions | it runs in one of the listed countries |
| hosts | it is one of the named hosts |
| require parameters | it accepts every parameter the request sends |

A host that has not stated a value for a constrained property does not qualify. A fallback that reaches several hosts through one route sends the constraints with the request, so only hosts that meet them serve it. A region constraint excludes such a route, because it makes no promise about where a host runs.

An endpoint can also be set to `primary_only`. It then never falls back; when the dedicated deployment cannot serve, you get the error.

For one request you can narrow the routes but never widen them. `"provider": {"allow_fallbacks": false}` keeps the request on your dedicated capacity. On open endpoints, `provider.only` and `provider.ignore` filter by provider name, as on any request. Closed endpoints select their providers and hosts through the contract configuration: `provider.only`, `provider.ignore`, `provider.order`, `provider.upstream` and `provider.require_parameters` return `400 invalid_request`, including empty selectors. Omit those fields on closed endpoints. `allow_fallbacks` remains supported.

## Closed endpoints

An endpoint is open unless your contract says otherwise: like any request, it tells you which deployment and host of the shared pool served you, and your dedicated capacity shows as `routerplus`. A closed endpoint never names a deployment or a host. Every surface you can read names RouterPlus as the provider, whichever route served the request, and apart from the route role nothing you can read differs between the routes:

| Surface | On a closed endpoint |
|---|---|
| `x-tm-provider` | `routerplus` |
| `x-tm-served-by`, `x-tm-upstream-model`, `x-tm-upstream-status`, `x-tm-dropped-params`, `x-tm-affinity-wait-ms` | absent |
| `x-tm-dedicated-endpoint`, `x-tm-route-role` | as on an open endpoint |
| `x-tm-remaining-rpm`, `x-tm-remaining-tpm` | counted against the endpoint's own limits only |
| Chat completion body | a fixed set of fields, below; `model` is the endpoint id |
| Stream chunks | the same fields in each chunk; a provider's SSE comment lines are dropped; a provider's error in the stream ends it with the gateway's own error event, and nothing follows that event |
| Error messages | they name `routerplus`, never a deployment; a provider's own error body is never relayed |
| Error statuses | a provider's 401, 402, 403 or 404 is a retryable 503 `upstream_error`; a `retry-after` is the gateway's own, never a provider's |
| Limits | a refusal because the capacity behind one route is full has `x-tm-limit-scope: endpoint` and `x-tm-limit-id: endpoint:<endpoint id>` |
| `provider.only`, `provider.ignore`, `provider.order`, `provider.upstream` | refused with 400 `invalid_request` |
| [`/v1/generation`](/docs/api-usage) | `deployment` is `routerplus` and `model` the endpoint id; `route_context` keeps only `dedicated` (the endpoint, its revision and the role); `admission_context` is null, `dropped_params` is empty, and admission events carry no pool scope id |
| [`/v1/usage`](/docs/api-usage) | `deployment` is `routerplus` and `model` the endpoint id |
| Console: Overview, Usage, Logs and a request's details | RouterPlus as the provider, no host, the endpoint id as the model, and "contract rate" as the price |
| The playground's receipt | served by RouterPlus |

The records stay masked once an endpoint has been closed: what it served before it closed, and after it is opened again, still reads as above.

The fields of a closed chat completion:

| Level | Fields |
|---|---|
| Top level | `id` (`chatcmpl-` and the request id without dashes), `object`, `created` (when the gateway received your request), `model`, `choices`, `usage` |
| A choice | `index`, `message` (not streamed) or `delta` (streamed), `finish_reason` (`stop`, `length`, `tool_calls`, `content_filter` or `function_call`; null until a stream ends), `logprobs` (only when you ask for them) |
| `message` or `delta` | `role`, `content`, `tool_calls` (`id`, `type`, `function.name`, `function.arguments`, and `index` in a stream), `function_call`, `refusal`, `reasoning_content` |
| `usage` | `prompt_tokens`, `completion_tokens`, `total_tokens`, `prompt_tokens_details.cached_tokens` and `completion_tokens_details.reasoning_tokens` (both always present, 0 when none), `cost` |

Every other field is dropped, including any field a provider adds later, and a field without a value is left out (a message keeps `content`: null beside tool calls). A model's thinking arrives as `reasoning_content` on every route. A tool call's id is the gateway's own, `call_` and 24 hex characters, the same on every chunk of the call and as the `tool_use` id on `/v1/messages`. In a stream, the first delta of a choice carries the role, the first delta of a tool call its id, type and name, and the finish and the usage each come in a chunk of their own.

## Prices

| Attempt | Price |
|---|---|
| On your dedicated capacity | your contract rate card, per million tokens |
| On a fallback | the model's list price, as for any other request; or your rate card, when your contract sets one price on every route |

A closed endpoint always bills your rate card on every route: a second price would tell you which route served. A fallback is tried only when its list price is known and at or below the endpoint's maximum price, if one is set, whichever price you then pay.

In `/v1/generation`, an attempt billed at your rate card has `inference_price_source` `customer_rate` and a `price_snapshot_id` of the form `dedicated:<endpoint id>:r<revision>`, so every charge traces back to the rate card that produced it. `usage.cost` on the response is the full debit across all attempts, as on every request.

## Limits

For the endpoint's traffic, its own limits replace the default organization, workspace, principal and key limits:

| Limit | Counted for |
|---|---|
| requests per minute | the endpoint |
| requests per second (optional): a count per second, or a paced rate | the endpoint |
| tokens per minute | the endpoint |
| concurrent requests | the endpoint |

A limit your organization set itself on [`/console/limits`](https://app.routerplus.com/console/limits) still applies on top. You can lower what one key or workspace may use; you cannot raise the endpoint's limits.

A refusal is the usual 429 `rate_limit` (see [Rate limits & spend caps](/docs/limits)), with `x-tm-limit-scope: endpoint`:

| Case | `x-tm-limit-kind` | `x-tm-limit-id` |
|---|---|---|
| Over the endpoint's RPM, RPS, TPM or concurrency | `rpm`, `rps`, `tpm` or `concurrency` | `endpoint:<endpoint id>` |
| Above a paced endpoint's hard rate | `rps` | `endpoint:<endpoint id>` |
| A paced endpoint's wait would be longer than its bound | `rps_queue` | `endpoint:<endpoint id>` |
| The dedicated capacity is full and the endpoint does not spill over to the shared pool | the measure that is full | `endpoint:<endpoint id>` |
| A closed endpoint: the capacity behind every route it could take is full | the measure that is full | `endpoint:<endpoint id>` |
| Fallback traffic is over the endpoint's fallback cap | `rpm`, `tpm` or `concurrency` | `endpoint-fallback:<endpoint id>` |

`GET /v1/limits?model=<your-namespace>/<name>` returns your key's effective limits on the endpoint and a `dedicated` object: the endpoint, its revision, the model, the output default, the fallback caps and whether it spills over.

If the endpoint's outage mode is `closed`, requests are refused while the shared counter store is unreachable. In the default `bounded_open` mode they are admitted under the endpoint's own limits, divided among the gateways that serve it, for up to 10 minutes; see [Limits and capacity](/docs/admission).

### The paced limit

A limit that counts each calendar second refuses requests in many seconds when traffic arrives at random, even at an average below the limit. A paced limit queues a short burst instead. It has three settings: the paced rate, a hard rate above it, and the longest wait.

- Requests leave admission no faster than the paced rate. A request that must wait waits in the gateway. `x-tm-queue-ms` on the response says how long, in milliseconds.
- A request whose wait would be longer than the bound gets 429, `x-tm-limit-kind: rps_queue`. Its `retry-after` says when its wait would fit.
- The hard rate refuses a request at once, `x-tm-limit-kind: rps`. It allows a burst as long as the bound (at least one second), so it never refuses a request the queue would hold: with a bound of a second or more, a full queue is what refuses.
- The wait comes before the request's own deadline starts, so it does not use the endpoint's total time.
- A waiting request holds no concurrency slot. Its place in the pace stays used even if you disconnect.

For example, with a paced rate of 6 requests a second, a hard rate of 6.5 and a 2 s bound:

| Average rate, arriving at random | Refused (429) | Typical wait |
|---|---|---|
| 4 a second | almost none | about 0.1 s |
| 5 a second | about 0.3 % | about 0.3 s |
| 6 a second | about 4 % | about 0.9 s |
| Steadily above 6 a second | the excess, once the queue is full | up to 2 s |

A queue cannot hold a steady excess: over time the endpoint passes its paced rate.

## Timeouts and defaults

| Setting | On other requests | On an endpoint |
|---|---|---|
| Attempts per request | up to 8 | 1 to 8; 3 unless set |
| Total time per request | 450 s | 1 to 600 s |
| Wait for a streamed attempt's response headers before the next route | 20 s | 1 to 120 s |
| Time one attempt may take on a request that is not streamed, before the next route | 450 s | 1 to 600 s, below the total |
| `max_tokens` written into a request that sets none | 4,096 | set per endpoint |

A shorter header wait, or attempt limit for requests that are not streamed, makes failover faster when the dedicated deployment stalls. A stalled request still takes that limit plus the fallback's own time. The attempt limit applies only while another route remains: the last route may run for the rest of the total time.

## Edges

- `POST /v1/route` does not explain dedicated endpoints yet.
- Health circuits are shared by every request to a deployment. Two failures in a row on the dedicated deployment send all its traffic to the fallbacks for 30 seconds. Per-endpoint circuit settings are not available yet.
- We change an endpoint's configuration on request. Each change is a new revision, and it reaches every gateway within seconds.
