Console
Get started/Routing policies

Routing policies and Azure connections

Ordered connections, explicit funding fallback, and Azure.

/llms.txt

A saved routing policy says which connections and marketplace deployments a key may use, in what order, and who pays. Policies apply to orgs, workspaces and keys.

Create and bind a policy

Use the site origin and the same session cookie, JSON content type, and same-origin mutation protection as BYOK connection management. Inference keys cannot administer policies; viewers can read them.

POST /api/routing-policies creates an immutable policy revision:

json
{
  "models": ["your-exact-model-id"],
  "connections": ["FIRST_CONNECTION_UUID", "SECOND_CONNECTION_UUID"],
  "funding": "byok_only",
  "allow_fallbacks": true,
  "max_attempts": 4,
  "timeout_ms": 450000
}

The response is 201 with {"policy_id":"UUID"}. Connections must belong to your org. Bind an existing key with POST /api/routing-policies/{id}/bind, body {"key_id":"KEY_UUID"}. Bind an org policy with the same endpoint and {"scope":"org"}. Org policies apply to all keys in the org; key policies can narrow them. A key without its own policy or single connection binding uses the org policy. A legacy single-connection binding remains an additional restriction when an org policy exists. Policy binding replaces a key's prior single-connection binding; binding a single connection replaces its key policy.

GET /api/routing-policies lists the newest 100 revisions. GET by policy UUID returns one. Updates create a new revision and explicitly rebind it. There is no delete or implicit unbind-to-house operation. House-only cached decisions may take up to five seconds to observe a new binding. Existing dispatched calls may finish on their pinned version.

Funding, ordering and limits

FieldBehavior
models1–100 exact IDs; all applicable policies must allow the requested model
connectionsUp to eight ordered org-owned BYOK connection UUIDs
house_deploymentsUp to eight explicit marketplace deployment IDs
funding: "byok_only"Requires connections, forbids house deployments; default
funding: "house_only"Requires house deployments, forbids connections
funding: "byok_with_house_fallback"Requires both; house remains the final phase
allow_fallbacksDefaults true; false limits selection to the first eligible route before health filtering
max_attemptsInteger 1–8; default four
timeout_msGeneration/admission execution deadline, 1,000–600,000 ms; default 450,000
regions, zdr, data_collectionHard restrictions against operator-supplied evidence; missing evidence fails closed

Under a policy, automatic fallback stays within the first selected provider label. OpenAI and Azure can both be authorized and explicitly selected, but a failed OpenAI call does not automatically switch to Azure. Exact-model equivalence remains a separate decision; attest_equivalence:true is rejected. Azure deployment mappings are customer declarations.

The effective restrictions are the intersection of org policy, workspace policy, key policy/binding, the server catalog, and request restrictions. Empty intersections fail closed. Lower limits win. Ordering uses request connection preference, then provider preference, then saved order and a stable deployment ID tie-breaker. Marketplace supply remains after all BYOK candidates for mixed funding. An unavailable primary does not authorize a fallback when fallbacks are disabled. No retry crosses the existing streaming commit boundary or a content-policy denial.

The execution deadline covers admission and generation after planning, including active streams. Request upload, bounded policy/price reads, and mandatory journal settlement/drain are separate phases. Timing out never abandons required durable settlement. A late SQL reservation is released and never dispatched.

Request controls and explanations

The inference provider object accepts:

json
{
  "only": ["openai"],
  "ignore": [],
  "order": ["openai"],
  "connections": ["CONNECTION_UUID"],
  "connection_order": ["CONNECTION_UUID"],
  "funding": ["byok"],
  "allow_fallbacks": false,
  "require_parameters": true,
  "upstream": ["deepinfra/bf16"]
}

only, ignore and order name route labels: house anthropic, openai and openrouter, and a BYOK connection's profile (openai, anthropic, azure, bedrock). A route label keeps that meaning even when its route does not serve the model. Any other label is an OpenRouter host tag. When the openrouter route runs, the gateway sends the host tags to OpenRouter as its own order, only and ignore, with your allow_fallbacks (default true). On POST /v1/decisions the labels are the decision providers: typesafe and openrouter-decisions for Jev, openrouter-decisions-free for Mercury Decide, bespokelabs for Bespoke Nimble v3, workers-ai for Clef and Clef-flash, routerplus for Decider 2B and Kev 4B, perplexity-decisions for Perplexity Decider v1 27B, levanto for Sage and fastino for GLiNER-2.5-Decide. Clef's label is workers-ai, not cloudflare: cloudflare stays an OpenRouter host tag (OpenRouter's Cloudflare host runs some chat models), and on POST /v1/decisions it names no provider. In the same way, Perplexity Decider's label is perplexity-decisions, not perplexity, which stays OpenRouter's host tag for Perplexity's chat models. Route labels win over host tags with the same spelling: only: ["anthropic"] means the Anthropic route. A host tag in only keeps the openrouter route open. upstream names hosts behind an aggregator deployment (OpenRouter provider tags), up to eight, tried in that order with no fallback to other hosts; with upstream, the host tags from order, only and ignore do not go. The tags are OpenRouter provider slugs; the hosts that run a model are its served_by entries in /api/models.json. To keep a session's prompt cache on one host, see the Agent integration guide. require_parameters: true refuses a deployment that would drop one of your request fields instead of dispatching without it. regions, zdr and data_collection are accepted here too as hard restrictions.

These preferences never grant connections, models, providers, or funds that saved authority does not allow. Price/latency/throughput sorting, max-price controls, raw request keys and arbitrary destinations are unsupported and return 400. Customers cannot write the evidence used to satisfy region and privacy restrictions, and a restriction is not a residency or ZDR guarantee.

POST /v1/route on the gateway accepts {"model":"ID","provider":{...}} using the inference key. It returns the plan, catalog and policy IDs, the ordered candidates with their funding and pinned house price IDs, the exclusions with a reason each, the OpenRouter host tags it would send as openrouter_provider (null when there are none), and the signal time and source, with advisory: true. It performs no provider request, decryption or reservation. Results are advisory; health and authorization can change before actual dispatch.

GET /v1/models uses the same restrictions. An optional URL-encoded JSON provider query parameter narrows the listing, and output_modalities=text,image,video,decisions filters by kind. Token counting uses the same plan but requires an Anthropic transport; it sends one count request and does not perform inference fallback. GET /v1/generation?id=... includes each attempt's durable route_context and fallback cause. Policy inference returns x-tm-route-plan-id for correlation.

Free BYOK still requires no marketplace credit. House fallback reserves credit only when reached and charges under the existing pricing rules. Earlier house holds remain until the request finishes, so fallback admission is conservative while settlements are buffered. Each attempt pins its own funding/credential identity. BYOK external cost stays unknown; house prices come from the existing model-level registry and are frozen before dispatch.

Azure/Foundry API-key connections

Create a connection through POST /api/connections:

json
{
  "profile": "azure",
  "models": ["your-public-model-id"],
  "api_key": "YOUR_RESOURCE_API_KEY",
  "endpoint_config": {
    "resource": "your-resource-name",
    "host": "foundry",
    "api_version": "v1",
    "deployments": {"your-public-model-id": "your-deployment-name"}
  }
}

host is openai ({resource}.openai.azure.com) or foundry ({resource}.services.ai.azure.com); the server constructs the public hostname, and customers cannot submit URLs. Every declared model needs a deployment mapping. Endpoint config is immutable; changes require a new connection. Validate, then bind it directly or include it in a policy. Rotation and revocation follow the connection lifecycle.

The adapter uses /openai/v1/chat/completions?api-version=v1, the deployment name in the model field, and the native api-key header. Validation calls the model-list endpoint with no generation. These shapes follow Microsoft's chat reference, models reference, and v1 lifecycle guide.

Only API-key authentication and public v1 endpoints are implemented. Entra/workload identity, legacy dated APIs, preview versions, sovereign/private endpoints, and Azure capacity/residency qualification are not included. Validation confirms endpoint authentication, not deployment entitlement, exact-model equivalence, quota, or capacity.

Shared admission requires declared quota pools and gateway limits. Quota availability may skip a candidate inside this frozen authority; it never grants new providers or marketplace funding. Model listings/explanations do not reserve future capacity.

Workspace bindings use the identity API. Durable attempt evidence stores at most 32 exclusion details and excludedTotal for the full count; the catalog hash and full route explanation preserve context without letting a large catalog exceed journal bounds.

Markdown source for agents: /docs/routing-policies.md · index at /llms.txt