# List your models

Sell your inference through the marketplace: one model document, a published
conformance bar, pass-through pricing, and automatic payment on the schedule
you choose.

Two things to know up front:

- **Pricing is pass-through.** Buyers pay your listed price; we never mark it
  up. Your price is your price.
- **Routing is earned.** Traffic flows on live health and published metrics —
  the same numbers you can see — not on negotiation.

**To apply:** email **providers@evo-hq.com** with your model document attached.
Onboarding is white-glove today: a human plus the automated suite take it from
there, usually same-week.

## The model document

Everything we need is one JSON document — your endpoint, models, prices,
capacity, and data policy. Where OpenRouter polls a `/v1/models` endpoint you
host, we accept the same information **pushed** as a versioned manifest: you
send updates (new models, price changes) when you choose, and nothing changes
under you between updates.

Validate against the published JSON Schema before sending:
[`/docs/provider-manifest.schema.json`](/docs/provider-manifest.schema.json).

```json
{
  "manifest_version": "1",
  "provider": {
    "id": "acme",
    "name": "Acme Inference",
    "privacy_policy_url": "https://acme.example/privacy",
    "terms_of_service_url": "https://acme.example/terms",
    "status_page_url": "https://status.acme.example",
    "prompt_logging": "none"
  },
  "endpoint": {
    "base_url": "https://api.acme.example/v1",
    "dialect": "openai",
    "api_key_env": "ACME_API_KEY"
  },
  "settlement": {
    "mode": "arrears",
    "interval": "weekly",
    "billing_contact": "billing@acme.example"
  },
  "models": [
    {
      "id": "acme-large-1",
      "display_name": "Acme Large 1",
      "context_length": 128000,
      "max_output_tokens": 16384,
      "streaming": true,
      "is_ready": false,
      "pricing": [
        { "type": "prompt", "unit": "token", "cost_usd_per_million": "0.90" },
        { "type": "completion", "unit": "token", "cost_usd_per_million": "3.60" }
      ],
      "capacity": [
        { "type": "completion", "unit": "token", "per": "minute", "value": 2000000 }
      ]
    }
  ]
}
```

### Identity

`id` is the **exact** identifier we send when calling your API — never an
alias. `display_name` is what buyers see. A listing is text (chat completions)
unless `output_modality` says `"image"`, `"video"` or `"decisions"`. An image listing is served
on `POST /v1/images/generations` only; it needs no `context_length` or
`streaming`, and must declare `max_output_tokens` (the per-image ceiling the
reservation is computed from) plus `supported_parameters` for `n` and a `size`
or `aspect_ratio` enum. A video listing is served on `POST /v1/videos`; its
`max_output_tokens` is the ceiling per second, and `seconds` must be declared
with a default. A decisions listing (a structured decision model that returns typed
answers to questions about a state) is served on `POST /v1/decisions`; it needs a
`context_length` but no `streaming`, and your endpoint must declare `decisions_path`, the
path we POST the request to. If your API picks the model by its URL, each decisions model
declares its own `decisions_path` instead (Cloudflare Workers AI: `/clef` and `/clef-flash`),
and then the endpoint needs none; the body still carries the model's `upstream_id` (or `id`)
as `model`. The model declares `decisions_wire`, the question format your
API speaks: `"systemone"` (TypeSafe's System One API, which Jev, Mercury Decide, Bespoke
Nimble v3, Clef, Clef-flash, Decider 2B, Kev 4B and Perplexity Decider v1 27B use; the default when the field is absent) or `"gliner"` (Fastino's GLiNER
schema, sent in an OpenAI chat-completions envelope, so `decisions_path` can be
`/chat/completions`). Buyers send questions in that format, and the gateway sends them to
you in it. A decisions listing is served only on `POST /v1/decisions`, even when its path is
a chat path: chat traffic never reaches it. Two listings of one model id must declare
the same wire. `decisions_wire` on a listing that is not `"decisions"` is refused.

If your API puts every answer and every error inside a transport envelope of its own,
declare it as `endpoint.response_envelope`, and the gateway opens it before it reads the
answer. Only a manifest of decisions models may declare it. The one value today is
`"cloudflare_v4"`, Cloudflare's REST envelope: a 200 is
`{"result": <answer>, "success": true, "errors": [], "messages": []}`, and an error is
`{"errors": [{"code": ..., "message": ...}], "success": false, "result": {}}`. A 200 whose
envelope does not say `success: true` is a failed answer, and nothing is billed for it.

Two more fields tell the gateway how your System One API reads a body. They are booleans on
a decisions model whose `decisions_wire` is `"systemone"` (or absent), and are refused on any
other listing. Absent means `false`.

- `decisions_state_per_question`: your API runs one prompt per question, each with the
  whole state, and bills the state once per question (Perplexity's decider does). The
  gateway then holds the state once per question when it reserves money and tokens for a
  request, so a request with many questions is held at what it will cost.
- `decisions_image_parts`: your API reads an object whose `type` is `"image_url"`, at any
  depth of the state or of a question, as an image. A decisions listing takes text and
  JSON, so the gateway refuses such a body with a typed 400 before any money is reserved,
  and it never reaches you.

A chat listing may declare `input_modalities`: what a message to the model may carry. The
list always has `"text"`, plus any of `"image"`, `"file"` (a document such as a PDF),
`"audio"` and `"video"`. Buyers see it on the model page and filter the catalog by it, and
the API lists it in `/v1/models`. The model page shows an input only when every listing of
that model declares it. The gateway sends a part only to a listing that declares its kind:
a part your listing does not declare never reaches you, and a part that no listing of the
model declares is refused with a typed 400. A declared part goes to you as sent when the
buyer calls the surface that speaks your dialect; across dialects the gateway refuses it
with a typed 400 and does not translate it. Absent means text only. An image, video or
decisions listing takes a text prompt, and the field is refused on it.

### Serving a model the marketplace already lists

If you serve a model another provider also lists (the marketplace id
`claude-sonnet-4-5`, say), list it under **that** id and set `upstream_id` to
the name your endpoint expects:

```json
{ "id": "claude-sonnet-4-5", "upstream_id": "anthropic/claude-sonnet-4.5", "...": "..." }
```

Requests to you carry `upstream_id`; conformance check C7 expects you to echo
it. Prices must be **identical** to the other providers of that id — the
catalog build rejects a mismatch (pass-through pricing means one price per
model). `provider.priority` orders providers of one model: `0` for a
first-party lab, `1` for an aggregator fallback; routing tries lower first.

### Endpoint

`dialect` is `"openai"` (Chat Completions) or `"anthropic"` (Messages) — one is
enough. Buyers reach your models from **both** marketplace surfaces regardless;
the gateway translates. `api_key_env` names the secret slot for the key you
issue us: it never appears in a manifest, page, or log.

### Pricing

Entries are `{type, unit, cost_usd_per_million}` with costs as decimal
**strings** in USD (never floats). Types: `prompt`, `cached_prompt`,
`cache_write`, `completion`, `internal_reasoning`; unit is `token`.

- **Declare what you charge.** A SKU you omit bills those tokens at your base
  `prompt`/`completion` rate — a missing SKU never means free, and never means
  a surprise for either side.
- **Price changes are append-only.** A new effective-dated row, never a
  rewrite. Every buyer request bills at the snapshot in force when it was
  dispatched, so your statement and their bill can never disagree about
  history.
- Flat per-token pricing only for now: conditional overrides, time-of-day
  windows, and per-request SKUs are not yet supported (see the OpenRouter
  mapping below).

### Capacity

Entries are `{type, unit, per, value}` — `request`, `prompt` or `completion`
limits per `minute`/`hour`/`day` window (`unit: "token"` is required for the two
token types). Identity is type + window; duplicates are rejected at validation.
Absent capacity means undeclared, not zero.

Capacity is information for us, not a setting. Validation checks it, but routing
and admission do not read it. The limits that admission enforces for your
deployment are a house pool that our operators set by hand from what you
declare and what you tell us (see [Limits and capacity](/docs/admission)). Above
that pool, your own 429 is the backstop: it pauses the pool for its
`retry-after` (1 to 60 seconds).

### Parameters

`supported_parameters` may be declared per model, and it is **enforced on the
image and video routes**: a value outside a declared enum is a 400 naming the
field. An image listing must declare `n` and a `size` or `aspect_ratio` enum, or
it fails validation; `quality` and `resolution` are optional but must be enums
when declared. On chat routes the gateway applies its own per-dialect
allowlists, and **every parameter it strips is recorded per-request** (visible
to the buyer in the `x-tm-dropped-params` header and their audit endpoint).
Declaring accurately now means chat enforcement picks it up when it lands.

### Data policy

`prompt_logging` is `"none"` or `"retained"`, published verbatim on your model
pages next to your privacy policy and terms URLs. Say the true thing — buyers
filter on it. Our own side is content-free by design: the marketplace never
stores prompts or responses, so your data policy is the only one in the path.

### Lifecycle

- **`is_ready: false`** stages a model: validated, conformance-testable, listed
  nowhere, routed never. Flip it when you are ready — a model proves itself
  before going live, not after.
- **`deprecation_date`** (ISO date): past it, the model leaves the catalog
  automatically.
- **`released`** (ISO date): the day the model came out. The Models page and the
  playground list models newest first after a few featured ones, so set it when you add
  a model: the lab's release date, or the day the model first listed publicly.
- Date-pinned requests (`acme-large-1-20260901`) resolve to your listed model
  for routing and billing, and the buyer's original id reaches your API
  verbatim — you decide whether to serve the snapshot.

## The conformance bar

Where OpenRouter runs unpublished "baseline tests," our bar is **published** —
seven checks, run against every model on both the stream and non-stream paths,
at listing time and on every catalog change. We share the report with you.

- **C1** — a non-stream completion returns provider-reported usage tokens.
  Usage is the billable record; no usage, no listing.
- **C2** — streamed responses are valid SSE, every frame parseable.
- **C3** — streams end with their dialect's terminal marker (`[DONE]` /
  `message_stop`).
- **C4** — usage tokens are reported in-stream.
- **C5** — an unknown model returns a parseable 4xx JSON error, not a 200 and
  not HTML.
- **C6** — a billed response actually contains assistant content. Usage
  counters without an answer is the one failure we will never pass.
- **C7** — the response reports the model that was asked for. Resolving an
  undated id to your own dated snapshot is fine; any other equivalence must be
  declared in `resolves_to`. Silently substituting a different model is
  disqualifying.

An **image** listing (`output_modality: "image"`) is non-stream only, so C2–C4 do not
apply; it runs C5 plus three image checks, each of which buys one real render at the
model's top declared quality:

- **I1** — a generation returns decodable base64 and provider-reported usage tokens.
- **I2** — the tokens billed for that render fit the declared per-image ceiling
  (`max_output_tokens`). An under-declared ceiling is a house loss on every request.
- **I3** — an undeclared size returns a parseable 4xx JSON error.
- **I4** (opt-in, `--image-ceiling`) — I2 at every declared size and at `size: auto`.

A **video** listing (`output_modality: "video"`) runs C5 plus three video checks:

- **V1** — a video job at the cheapest tier finishes, reports its cost and downloads as an MP4.
- **V2** — the declared per-second ceiling covers the reported cost.
- **V3** — an undeclared shape is refused, not silently substituted.

A **decisions** listing (`output_modality: "decisions"`) runs two checks:

- **D1** — one call with three questions comes back as typed answers, with usage. On the
  `systemone` wire: one Choice, one Score and one Noul question, and `usage.input_tokens`.
  On the `gliner` wire: one yes/no task and one three-way task in one schema (sent with
  `store: false`), the message content parsing to
  one JSON object with a labelled answer under each task name, and `prompt_tokens` in the
  usage. D1 calls each model at its own `decisions_path` when the model declares one, else
  at the endpoint's. With `endpoint.response_envelope: "cloudflare_v4"`, the 200 must say
  `success: true`, and D1 reads the answer and the usage inside `result`. When D1 gets
  another status, the report shows the first 300 characters of your body.
- **C5** — an unknown model yields a parseable 4xx error. FastAPI's `{"detail": ...}`
  envelope counts as a described error, and so does Cloudflare's
  `{"errors": [{"code": ..., "message": ...}]}` (or one such object in place of the list)
  when an entry has a message.

For every model, the report also lists any `Ratelimit`, `Ratelimit-Policy` or `Retry-After`
header your endpoint sent, or says that it sent none. These headers are reported, never
judged.

A provider that reports its own cost per render instead of tokens declares
`endpoint.billing: "reported_cost"`; the ledger then charges exactly the cost
reported, and the image and video checks judge that cost against the ceiling.

## Wire requirements

- HTTPS endpoint speaking OpenAI Chat Completions **or** Anthropic Messages,
  with streaming.
- Provider-reported usage tokens on both paths (C1/C4).
- Honest errors: return early 429s at capacity — queueing wrecks your published
  latency and your buyers' experience; 4xx JSON for bad requests.
- **Streaming:** response headers within 20 seconds or the attempt fails over;
  send SSE keep-alive comments through long pre-first-byte gaps, and stream
  tokens as soon as they exist. Non-stream requests are exempt — they get a
  long generation budget instead.
- After first output we never switch providers mid-answer — a stall is
  surfaced to the buyer as yours. Stream steadily.

## How uptime is counted

Exact accounting, same numbers routing uses:

```text
uptime = completed ÷ (completed + counted failures), trailing 24h
```

- **Counted against you:** provider 401/402/403, all 5xx, mid-stream deaths,
  incomplete streams.
- **Never counted:** buyer-caused requests — 400 and 413 (bad input), 429
  (their rate limit), and client cancellations mid-stream.
- **Published only after 100+ requests.** Below that the pages say `n<100` —
  we never print a fake 100%, yours or anyone's.
- Failures open a per-deployment **health circuit**: cooldown, then exactly one
  half-open probe; recovery closes it. While everything for a model cools down
  we still send a last-resort attempt rather than strand buyers. Content-policy
  refusals are never retried and never failed over.

## Performance

We publish **TTFB p50/p95** (from merged t-digests) per model — the same
figures the router sees. Two behaviors dominate them: return early 429s instead
of queueing, and start streaming immediately.

## How you get paid

Pick a settlement mode in your manifest — the equivalent of OpenRouter's
"auto top-up or invoicing" requirement, both directions first-class:

- **`prepaid`** — we deposit with you ahead of traffic; usage draws it down,
  and we monitor balance against burn so routing never hits a dry account.
- **`arrears`** — we route first and pay on your interval: `half_daily`,
  `daily`, `weekly`, or `monthly`.

Your statement is generated from the **same append-only ledger that bills
buyers** — token counts and USD per period, auditable to the individual
attempt, with per-attempt byte counts recorded as a tokenizer-independent bound
for disputes.

## What you get

- Buyer traffic from both wire surfaces the day your models go live.
- A public model page per model: prices, context, your data-policy flag, live
  honest metrics.
- Per-period settlement statements that reconcile to the token.
- Degraded providers get circuits and time to recover, not silent delisting.

## Coming from OpenRouter?

The concepts map directly; here is the translation:

| OpenRouter | Here |
|---|---|
| Polled `/v1/models` model documents | Pushed manifest (this page) — same information, updated when you choose |
| `is_ready`, `deprecation_date` | Same names, same semantics |
| Baseline tests (unpublished) | C1–C7, I1–I4 and V1–V3, published above, report shared with you |
| Uptime after 100+ requests, user errors excluded | Same — formula published above |
| Data-policy / may-train disclosure | `prompt_logging` + your policy URLs, shown to buyers |
| Auto top-up or invoicing | `prepaid` or `arrears` with ledger-audited statements |
| Conditional pricing, time-of-day windows, `discount_to_user`, `:free` variants | Not yet — flat per-token pricing only |
| Multimodal listings | Image and video generation, and decision models, via `output_modality`; image, file, audio and video inputs to chat models via `input_modalities` (passed through in your own dialect, never translated; a kind you do not declare never reaches you) |
| Provider dashboard | Not yet — statements and conformance reports by email while the console is built |

## Apply

Email **providers@evo-hq.com** with your manifest attached (validate it against
[the schema](/docs/provider-manifest.schema.json) first). We run the suite
against your endpoint, share the report, stage your models, and flip them live
together.
