Console
For providers/List your models

List your models

One manifest, a published conformance bar for text, image and video, paid prepaid or in arrears.

/llms.txt

Sell your inference through the marketplace: one model document, a published conformance bar, pass-through pricing, and automatic payment on the schedule you choose.

Two things to know up front:

  • Pricing is pass-through. Buyers pay your listed price; we never mark it up. Your price is your price.
  • Routing is earned. Traffic flows on live health and published metrics — the same numbers you can see — not on negotiation.

To apply: email [email protected] with your model document attached. Onboarding is white-glove today: a human plus the automated suite take it from there, usually same-week.

The model document

Everything we need is one JSON document — your endpoint, models, prices, capacity, and data policy. Where OpenRouter polls a /v1/models endpoint you host, we accept the same information pushed as a versioned manifest: you send updates (new models, price changes) when you choose, and nothing changes under you between updates.

Validate against the published JSON Schema before sending: /docs/provider-manifest.schema.json.

json
{
  "manifest_version": "1",
  "provider": {
    "id": "acme",
    "name": "Acme Inference",
    "privacy_policy_url": "https://acme.example/privacy",
    "terms_of_service_url": "https://acme.example/terms",
    "status_page_url": "https://status.acme.example",
    "prompt_logging": "none"
  },
  "endpoint": {
    "base_url": "https://api.acme.example/v1",
    "dialect": "openai",
    "api_key_env": "ACME_API_KEY"
  },
  "settlement": {
    "mode": "arrears",
    "interval": "weekly",
    "billing_contact": "[email protected]"
  },
  "models": [
    {
      "id": "acme-large-1",
      "display_name": "Acme Large 1",
      "context_length": 128000,
      "max_output_tokens": 16384,
      "streaming": true,
      "is_ready": false,
      "pricing": [
        { "type": "prompt", "unit": "token", "cost_usd_per_million": "0.90" },
        { "type": "completion", "unit": "token", "cost_usd_per_million": "3.60" }
      ],
      "capacity": [
        { "type": "completion", "unit": "token", "per": "minute", "value": 2000000 }
      ]
    }
  ]
}

Identity

id is the exact identifier we send when calling your API — never an alias. display_name is what buyers see. A listing is text (chat completions) unless output_modality says "image", "video" or "decisions". An image listing is served on POST /v1/images/generations only; it needs no context_length or streaming, and must declare max_output_tokens (the per-image ceiling the reservation is computed from) plus supported_parameters for n and a size or aspect_ratio enum. A video listing is served on POST /v1/videos; its max_output_tokens is the ceiling per second, and seconds must be declared with a default. A decisions listing (a structured decision model that returns typed answers to questions about a state) is served on POST /v1/decisions; it needs a context_length but no streaming, and your endpoint must declare decisions_path, the path we POST the request to. If your API picks the model by its URL, each decisions model declares its own decisions_path instead (Cloudflare Workers AI: /clef and /clef-flash), and then the endpoint needs none; the body still carries the model's upstream_id (or id) as model. The model declares decisions_wire, the question format your API speaks: "systemone" (TypeSafe's System One API, which Jev, Mercury Decide, Bespoke Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B and Perplexity Decider v1 27B use; the default when the field is absent) or "gliner" (Fastino's GLiNER schema, sent in an OpenAI chat-completions envelope, so decisions_path can be /chat/completions). Buyers send questions in that format, and the gateway sends them to you in it. A decisions listing is served only on POST /v1/decisions, even when its path is a chat path: chat traffic never reaches it. Two listings of one model id must declare the same wire. decisions_wire on a listing that is not "decisions" is refused.

If your API puts every answer and every error inside a transport envelope of its own, declare it as endpoint.response_envelope, and the gateway opens it before it reads the answer. Only a manifest of decisions models may declare it. The one value today is "cloudflare_v4", Cloudflare's REST envelope: a 200 is {"result": <answer>, "success": true, "errors": [], "messages": []}, and an error is {"errors": [{"code": ..., "message": ...}], "success": false, "result": {}}. A 200 whose envelope does not say success: true is a failed answer, and nothing is billed for it.

Two more fields tell the gateway how your System One API reads a body. They are booleans on a decisions model whose decisions_wire is "systemone" (or absent), and are refused on any other listing. Absent means false.

  • decisions_state_per_question: your API runs one prompt per question, each with the whole state, and bills the state once per question (Perplexity's decider does). The gateway then holds the state once per question when it reserves money and tokens for a request, so a request with many questions is held at what it will cost.
  • decisions_image_parts: your API reads an object whose type is "image_url", at any depth of the state or of a question, as an image. A decisions listing takes text and JSON, so the gateway refuses such a body with a typed 400 before any money is reserved, and it never reaches you.

A chat listing may declare input_modalities: what a message to the model may carry. The list always has "text", plus any of "image", "file" (a document such as a PDF), "audio" and "video". Buyers see it on the model page and filter the catalog by it, and the API lists it in /v1/models. The model page shows an input only when every listing of that model declares it. The gateway sends a part only to a listing that declares its kind: a part your listing does not declare never reaches you, and a part that no listing of the model declares is refused with a typed 400. A declared part goes to you as sent when the buyer calls the surface that speaks your dialect; across dialects the gateway refuses it with a typed 400 and does not translate it. Absent means text only. An image, video or decisions listing takes a text prompt, and the field is refused on it.

Serving a model the marketplace already lists

If you serve a model another provider also lists (the marketplace id claude-sonnet-4-5, say), list it under that id and set upstream_id to the name your endpoint expects:

json
{ "id": "claude-sonnet-4-5", "upstream_id": "anthropic/claude-sonnet-4.5", "...": "..." }

Requests to you carry upstream_id; conformance check C7 expects you to echo it. Prices must be identical to the other providers of that id — the catalog build rejects a mismatch (pass-through pricing means one price per model). provider.priority orders providers of one model: 0 for a first-party lab, 1 for an aggregator fallback; routing tries lower first.

Endpoint

dialect is "openai" (Chat Completions) or "anthropic" (Messages) — one is enough. Buyers reach your models from both marketplace surfaces regardless; the gateway translates. api_key_env names the secret slot for the key you issue us: it never appears in a manifest, page, or log.

Pricing

Entries are {type, unit, cost_usd_per_million} with costs as decimal strings in USD (never floats). Types: prompt, cached_prompt, cache_write, completion, internal_reasoning; unit is token.

  • Declare what you charge. A SKU you omit bills those tokens at your base prompt/completion rate — a missing SKU never means free, and never means a surprise for either side.
  • Price changes are append-only. A new effective-dated row, never a rewrite. Every buyer request bills at the snapshot in force when it was dispatched, so your statement and their bill can never disagree about history.
  • Flat per-token pricing only for now: conditional overrides, time-of-day windows, and per-request SKUs are not yet supported (see the OpenRouter mapping below).

Capacity

Entries are {type, unit, per, value} — request, prompt or completion limits per minute/hour/day window (unit: "token" is required for the two token types). Identity is type + window; duplicates are rejected at validation. Absent capacity means undeclared, not zero.

Capacity is information for us, not a setting. Validation checks it, but routing and admission do not read it. The limits that admission enforces for your deployment are a house pool that our operators set by hand from what you declare and what you tell us (see Limits and capacity). Above that pool, your own 429 is the backstop: it pauses the pool for its retry-after (1 to 60 seconds).

Parameters

supported_parameters may be declared per model, and it is enforced on the image and video routes: a value outside a declared enum is a 400 naming the field. An image listing must declare n and a size or aspect_ratio enum, or it fails validation; quality and resolution are optional but must be enums when declared. On chat routes the gateway applies its own per-dialect allowlists, and every parameter it strips is recorded per-request (visible to the buyer in the x-tm-dropped-params header and their audit endpoint). Declaring accurately now means chat enforcement picks it up when it lands.

Data policy

prompt_logging is "none" or "retained", published verbatim on your model pages next to your privacy policy and terms URLs. Say the true thing — buyers filter on it. Our own side is content-free by design: the marketplace never stores prompts or responses, so your data policy is the only one in the path.

Lifecycle

  • is_ready: false stages a model: validated, conformance-testable, listed nowhere, routed never. Flip it when you are ready — a model proves itself before going live, not after.
  • deprecation_date (ISO date): past it, the model leaves the catalog automatically.
  • released (ISO date): the day the model came out. The Models page and the playground list models newest first after a few featured ones, so set it when you add a model: the lab's release date, or the day the model first listed publicly.
  • Date-pinned requests (acme-large-1-20260901) resolve to your listed model for routing and billing, and the buyer's original id reaches your API verbatim — you decide whether to serve the snapshot.

The conformance bar

Where OpenRouter runs unpublished "baseline tests," our bar is published — seven checks, run against every model on both the stream and non-stream paths, at listing time and on every catalog change. We share the report with you.

  • C1 — a non-stream completion returns provider-reported usage tokens. Usage is the billable record; no usage, no listing.
  • C2 — streamed responses are valid SSE, every frame parseable.
  • C3 — streams end with their dialect's terminal marker ([DONE] / message_stop).
  • C4 — usage tokens are reported in-stream.
  • C5 — an unknown model returns a parseable 4xx JSON error, not a 200 and not HTML.
  • C6 — a billed response actually contains assistant content. Usage counters without an answer is the one failure we will never pass.
  • C7 — the response reports the model that was asked for. Resolving an undated id to your own dated snapshot is fine; any other equivalence must be declared in resolves_to. Silently substituting a different model is disqualifying.

An image listing (output_modality: "image") is non-stream only, so C2–C4 do not apply; it runs C5 plus three image checks, each of which buys one real render at the model's top declared quality:

  • I1 — a generation returns decodable base64 and provider-reported usage tokens.
  • I2 — the tokens billed for that render fit the declared per-image ceiling (max_output_tokens). An under-declared ceiling is a house loss on every request.
  • I3 — an undeclared size returns a parseable 4xx JSON error.
  • I4 (opt-in, --image-ceiling) — I2 at every declared size and at size: auto.

A video listing (output_modality: "video") runs C5 plus three video checks:

  • V1 — a video job at the cheapest tier finishes, reports its cost and downloads as an MP4.
  • V2 — the declared per-second ceiling covers the reported cost.
  • V3 — an undeclared shape is refused, not silently substituted.

A decisions listing (output_modality: "decisions") runs two checks:

  • D1 — one call with three questions comes back as typed answers, with usage. On the systemone wire: one Choice, one Score and one Noul question, and usage.input_tokens. On the gliner wire: one yes/no task and one three-way task in one schema (sent with store: false), the message content parsing to one JSON object with a labelled answer under each task name, and prompt_tokens in the usage. D1 calls each model at its own decisions_path when the model declares one, else at the endpoint's. With endpoint.response_envelope: "cloudflare_v4", the 200 must say success: true, and D1 reads the answer and the usage inside result. When D1 gets another status, the report shows the first 300 characters of your body.
  • C5 — an unknown model yields a parseable 4xx error. FastAPI's {"detail": ...} envelope counts as a described error, and so does Cloudflare's {"errors": [{"code": ..., "message": ...}]} (or one such object in place of the list) when an entry has a message.

For every model, the report also lists any Ratelimit, Ratelimit-Policy or Retry-After header your endpoint sent, or says that it sent none. These headers are reported, never judged.

A provider that reports its own cost per render instead of tokens declares endpoint.billing: "reported_cost"; the ledger then charges exactly the cost reported, and the image and video checks judge that cost against the ceiling.

Wire requirements

  • HTTPS endpoint speaking OpenAI Chat Completions or Anthropic Messages, with streaming.
  • Provider-reported usage tokens on both paths (C1/C4).
  • Honest errors: return early 429s at capacity — queueing wrecks your published latency and your buyers' experience; 4xx JSON for bad requests.
  • Streaming: response headers within 20 seconds or the attempt fails over; send SSE keep-alive comments through long pre-first-byte gaps, and stream tokens as soon as they exist. Non-stream requests are exempt — they get a long generation budget instead.
  • After first output we never switch providers mid-answer — a stall is surfaced to the buyer as yours. Stream steadily.

How uptime is counted

Exact accounting, same numbers routing uses:

uptime = completed ÷ (completed + counted failures), trailing 24h
  • Counted against you: provider 401/402/403, all 5xx, mid-stream deaths, incomplete streams.
  • Never counted: buyer-caused requests — 400 and 413 (bad input), 429 (their rate limit), and client cancellations mid-stream.
  • Published only after 100+ requests. Below that the pages say n<100 — we never print a fake 100%, yours or anyone's.
  • Failures open a per-deployment health circuit: cooldown, then exactly one half-open probe; recovery closes it. While everything for a model cools down we still send a last-resort attempt rather than strand buyers. Content-policy refusals are never retried and never failed over.

Performance

We publish TTFB p50/p95 (from merged t-digests) per model — the same figures the router sees. Two behaviors dominate them: return early 429s instead of queueing, and start streaming immediately.

How you get paid

Pick a settlement mode in your manifest — the equivalent of OpenRouter's "auto top-up or invoicing" requirement, both directions first-class:

  • prepaid — we deposit with you ahead of traffic; usage draws it down, and we monitor balance against burn so routing never hits a dry account.
  • arrears — we route first and pay on your interval: half_daily, daily, weekly, or monthly.

Your statement is generated from the same append-only ledger that bills buyers — token counts and USD per period, auditable to the individual attempt, with per-attempt byte counts recorded as a tokenizer-independent bound for disputes.

What you get

  • Buyer traffic from both wire surfaces the day your models go live.
  • A public model page per model: prices, context, your data-policy flag, live honest metrics.
  • Per-period settlement statements that reconcile to the token.
  • Degraded providers get circuits and time to recover, not silent delisting.

Coming from OpenRouter?

The concepts map directly; here is the translation:

OpenRouterHere
Polled /v1/models model documentsPushed manifest (this page) — same information, updated when you choose
is_ready, deprecation_dateSame names, same semantics
Baseline tests (unpublished)C1–C7, I1–I4 and V1–V3, published above, report shared with you
Uptime after 100+ requests, user errors excludedSame — formula published above
Data-policy / may-train disclosureprompt_logging + your policy URLs, shown to buyers
Auto top-up or invoicingprepaid or arrears with ledger-audited statements
Conditional pricing, time-of-day windows, discount_to_user, :free variantsNot yet — flat per-token pricing only
Multimodal listingsImage and video generation, and decision models, via output_modality; image, file, audio and video inputs to chat models via input_modalities (passed through in your own dialect, never translated; a kind you do not declare never reaches you)
Provider dashboardNot yet — statements and conformance reports by email while the console is built

Apply

Email [email protected] with your manifest attached (validate it against the schema first). We run the suite against your endpoint, share the report, stage your models, and flip them live together.

Markdown source for agents: /docs/for-providers.md · index at /llms.txt