# Routing policies and Azure connections

A saved routing policy says which connections and marketplace deployments a key may use,
in what order, and who pays. Policies apply to orgs, workspaces and keys.

## Create and bind a policy

Use the site origin and the same session cookie, JSON content type, and same-origin
mutation protection as [BYOK connection management](/docs/byok). Inference keys cannot
administer policies; viewers can read them.

`POST /api/routing-policies` creates an immutable policy revision:

```json
{
  "models": ["your-exact-model-id"],
  "connections": ["FIRST_CONNECTION_UUID", "SECOND_CONNECTION_UUID"],
  "funding": "byok_only",
  "allow_fallbacks": true,
  "max_attempts": 4,
  "timeout_ms": 450000
}
```

The response is `201` with `{"policy_id":"UUID"}`. Connections must belong to your org. Bind an
existing key with `POST /api/routing-policies/{id}/bind`, body `{"key_id":"KEY_UUID"}`.
Bind an org policy with the same endpoint and `{"scope":"org"}`. Org policies apply to
all keys in the org; key policies can narrow them. A key without its own policy or single
connection binding uses the org policy. A legacy single-connection binding remains an
additional restriction when an org policy exists. Policy binding replaces a key's prior
single-connection binding; binding a single connection replaces its key policy.

`GET /api/routing-policies` lists the newest 100 revisions. GET by policy UUID returns one.
Updates create a new revision and explicitly rebind it. There is no delete or implicit
unbind-to-house operation. House-only cached decisions may take up to five seconds to
observe a new binding. Existing dispatched calls may finish on their pinned version.

## Funding, ordering and limits

| Field | Behavior |
|---|---|
| `models` | 1–100 exact IDs; all applicable policies must allow the requested model |
| `connections` | Up to eight ordered org-owned BYOK connection UUIDs |
| `house_deployments` | Up to eight explicit marketplace deployment IDs |
| `funding: "byok_only"` | Requires connections, forbids house deployments; default |
| `funding: "house_only"` | Requires house deployments, forbids connections |
| `funding: "byok_with_house_fallback"` | Requires both; house remains the final phase |
| `allow_fallbacks` | Defaults true; false limits selection to the first eligible route before health filtering |
| `max_attempts` | Integer 1–8; default four |
| `timeout_ms` | Generation/admission execution deadline, 1,000–600,000 ms; default 450,000 |
| `regions`, `zdr`, `data_collection` | Hard restrictions against operator-supplied evidence; missing evidence fails closed |

Under a policy, automatic fallback stays within the first selected **provider label**. OpenAI
and Azure can both be authorized and explicitly selected, but a failed OpenAI call does not
automatically switch to Azure. Exact-model equivalence remains a separate decision;
`attest_equivalence:true` is rejected. Azure deployment mappings are customer declarations.

The effective restrictions are the intersection of org policy, workspace policy, key policy/binding, the
server catalog, and request restrictions. Empty intersections fail closed. Lower limits win.
Ordering uses request connection preference, then provider preference, then saved order and
a stable deployment ID tie-breaker. Marketplace supply remains after all BYOK candidates
for mixed funding. An unavailable primary does not authorize a fallback when fallbacks are
disabled. No retry crosses the existing streaming commit boundary or a content-policy denial.

The execution deadline covers admission and generation after planning, including active
streams. Request upload, bounded policy/price reads, and mandatory journal settlement/drain
are separate phases. Timing out never abandons required durable settlement. A late SQL
reservation is released and never dispatched.

## Request controls and explanations

The inference `provider` object accepts:

```json
{
  "only": ["openai"],
  "ignore": [],
  "order": ["openai"],
  "connections": ["CONNECTION_UUID"],
  "connection_order": ["CONNECTION_UUID"],
  "funding": ["byok"],
  "allow_fallbacks": false,
  "require_parameters": true,
  "upstream": ["deepinfra/bf16"]
}
```

`only`, `ignore` and `order` name route labels: house `anthropic`, `openai` and
`openrouter`, and a BYOK connection's profile (`openai`, `anthropic`, `azure`, `bedrock`). A
route label keeps that meaning even when its route does not serve the model. Any other
label is an OpenRouter host tag. When the `openrouter`
route runs, the gateway sends the host tags to OpenRouter as its own `order`, `only` and
`ignore`, with your `allow_fallbacks` (default true). On `POST /v1/decisions` the labels are
the decision providers: `typesafe` and `openrouter-decisions` for Jev,
`openrouter-decisions-free` for Mercury Decide, `bespokelabs` for Bespoke Nimble v3,
`workers-ai` for Clef and Clef-flash, `routerplus` for Decider 2B and Kev 4B,
`perplexity-decisions` for Perplexity Decider v1 27B, `levanto` for Sage and `fastino` for
GLiNER-2.5-Decide.
Clef's label is `workers-ai`, not `cloudflare`: `cloudflare` stays an OpenRouter host tag
(OpenRouter's Cloudflare host runs some chat models), and on `POST /v1/decisions` it names
no provider. In the same way, Perplexity Decider's label is `perplexity-decisions`, not
`perplexity`, which stays OpenRouter's host tag for Perplexity's chat models. Route labels win over host tags with the
same spelling: `only: ["anthropic"]` means the Anthropic route. A host tag in `only` keeps the
`openrouter` route open. `upstream` names hosts behind an aggregator deployment (OpenRouter
provider tags), up to eight, tried in that order with no fallback to other hosts; with
`upstream`, the host tags from `order`, `only` and `ignore` do not go. The tags are OpenRouter provider slugs; the hosts that run a model
are its `served_by` entries in [`/api/models.json`](https://app.routerplus.com/api/models.json). To keep a
session's prompt cache on one host, see the [Agent integration guide](agent-integration.md). `require_parameters: true`
refuses a deployment that would drop one of your request fields instead of dispatching
without it. `regions`, `zdr` and `data_collection` are accepted here too as hard
restrictions.

These preferences never grant connections, models, providers, or funds that saved authority
does not allow. Price/latency/throughput sorting, max-price controls, raw request keys and
arbitrary destinations are unsupported and return 400. Customers cannot write the evidence
used to satisfy region and privacy restrictions, and a restriction is not a residency or
ZDR guarantee.

`POST /v1/route` on the gateway accepts `{"model":"ID","provider":{...}}` using the
inference key. It returns the plan, catalog and policy IDs, the ordered candidates with
their funding and pinned house price IDs, the exclusions with a reason each, the OpenRouter
host tags it would send as `openrouter_provider` (`null` when there are none), and the
signal time and source, with `advisory: true`. It performs no provider request, decryption or
reservation. Results are advisory; health and authorization can change before actual dispatch.

`GET /v1/models` uses the same restrictions. An optional URL-encoded JSON `provider` query
parameter narrows the listing, and `output_modalities=text,image,video,decisions` filters by kind. Token counting uses the same plan but requires an Anthropic
transport; it sends one count request and does not perform inference fallback.
`GET /v1/generation?id=...` includes each attempt's durable `route_context` and fallback cause.
Policy inference returns `x-tm-route-plan-id` for correlation.

Free BYOK still requires no marketplace credit. House fallback reserves credit only when
reached and charges under the existing pricing rules. Earlier house holds remain until the
request finishes, so fallback admission is conservative while settlements are buffered.
Each attempt pins its own funding/credential identity. BYOK external cost stays unknown;
house prices come from the existing model-level registry and are frozen before dispatch.

## Azure/Foundry API-key connections

Create a connection through `POST /api/connections`:

```json
{
  "profile": "azure",
  "models": ["your-public-model-id"],
  "api_key": "YOUR_RESOURCE_API_KEY",
  "endpoint_config": {
    "resource": "your-resource-name",
    "host": "foundry",
    "api_version": "v1",
    "deployments": {"your-public-model-id": "your-deployment-name"}
  }
}
```

`host` is `openai` (`{resource}.openai.azure.com`) or `foundry`
(`{resource}.services.ai.azure.com`); the server constructs the public hostname, and
customers cannot submit URLs. Every declared model needs a deployment mapping. Endpoint
config is immutable; changes require a new connection. Validate, then bind it directly or
include it in a policy. Rotation and revocation follow the connection lifecycle.

The adapter uses `/openai/v1/chat/completions?api-version=v1`, the deployment name in the
model field, and the native `api-key` header. Validation calls the model-list endpoint with
no generation. These shapes follow Microsoft's [chat reference](https://learn.microsoft.com/en-us/rest/api/microsoft-foundry/azureopenai/chat),
[models reference](https://learn.microsoft.com/en-us/rest/api/microsoft-foundry/azureopenai/models),
and [v1 lifecycle guide](https://learn.microsoft.com/en-us/azure/foundry/openai/api-version-lifecycle).

Only API-key authentication and public v1 endpoints are implemented. Entra/workload identity,
legacy dated APIs, preview versions, sovereign/private endpoints, and Azure capacity/residency
qualification are not included. Validation confirms endpoint authentication, not deployment
entitlement, exact-model equivalence, quota, or capacity.

Shared admission requires declared [quota pools and gateway limits](/docs/admission).
Quota availability may skip a candidate inside this frozen authority; it never grants new
providers or marketplace funding. Model listings/explanations do not reserve future capacity.

Workspace bindings use the [identity API](identity.md). Durable attempt evidence stores at
most 32 exclusion details and `excludedTotal` for the full count; the catalog hash and full
route explanation preserve context without letting a large catalog exceed journal bounds.
