# Playground

[`https://app.routerplus.com/playground`](https://app.routerplus.com/playground) is an interface over the same
catalog, routing, prices and ledger as the API: chat, typed decisions, images and video.
Sign in, pick a model, and each answer arrives with the provider that served it, the token
counts, and what it cost.

## What it does

- **Any catalog model.** The picker searches the catalog; `⌘K` / `Ctrl+K`
  opens it. It lists the models of the mode you are in first (chat models in Chat,
  decision models in Decision). Chat starts with GPT-6 Astra, Claude Opus 5.5, GLM 5.3
  Flash, DeepSeek V4.1 Flash and Hy4 Preview; then the other open-weight models; then
  GPT-6.1 Sol, GPT-6 Luna and Claude Sonnet 5.5; then the rest, each group newest first.
  Decision mode lists Jev 1.13, Bespoke Nimble v3, Clef, Clef-flash, Mercury Decide, Kev 4B,
  GLiNER-2.5-Decide, Decider 2B and Sage 1.2, in that order. The search matches a model's name,
  id, lab, the providers that run it, and its kind: type `decision` for every decision model,
  `image` or `video` for those, `free` for the free ones. Each row carries the lab's mark,
  its price — per million tokens, or for an image or video model the provider's own price
  and the most one image or second can cost — and a chat model's context length. Typing
  an id that is not listed offers it verbatim, which is how a dated snapshot of a catalog
  model is pinned.
- **Compare up to three chat models side by side** (Decision mode compares up to four). **Compare** adds a column. One prompt goes
  to every column at once, each streams into its own card, and each shows its own provider,
  tokens, latency and cost. Every column keeps its own thread: a follow-up sees that model's
  replies only, never another's. Removing a column leaves the rest untouched.
- **Sample prompts** under the composer send a real prompt in one click. The first is
  always the **strawberry test** (how many times "r" appears in "strawberry"), asked
  plainly. Each reply to it says **passed**, **failed** with the count the model gave, or
  that it gave no clear count. The others are drawn at random: a reasoning trap, a code
  task, structured JSON, a table, a six-word story and a few more. They make way as soon as
  the conversation starts.
- **Streaming replies** with a stop button, regenerate, a collapsible "thinking" section for
  reasoning models, and safe rendering of code blocks, lists, links and tables. A table
  is drawn only when the reply holds a real markdown table: a header row, then a divider
  row such as `|---|---|` with the same number of cells. Any other text with a pipe in it
  reads as it is. A wide table scrolls sideways inside its card.
- **Code and Parameters** sit in the composer, to the left of Send. **Code** shows the
  request the playground is about to make, as curl, Python or Node, so you can paste it
  into your own project. **Parameters** holds a system prompt, temperature, top_p and max
  output tokens, per conversation.
- **Cost on every reply:** the same `usage.cost` the API returns, plus time to first byte and
  an `audit` link to the request's ledger rows.
- **Organizations:** if you belong to several, the sidebar selector switches which one the
  playground bills.

## Tetris Royale

[Tetris Royale](https://app.routerplus.com/playground/tetris-royale) compares the eight current
System One decision models on the same block sequence. Open it from the Playground
mode bar; the arrow button at the top left goes back. It keeps the selected organization and uses a separate server-held managed
key named **Tetris Royale**; ordinary Playground keys retain their existing limit.
The server validates the board and match counters, then constructs the decision prompt
and legal choices itself. The game relay does not accept custom prompts or questions.
Every model request follows the organization's normal credits, limits, billing and
provider routing. No API key reaches the browser.

A match defaults to 200 pieces; Settings offers 20 up to a thousand. Each model plays independently: a slow or rate-limited
model does not hold up the others. Transient 429s cool down only that model and retry
up to three times, respecting Retry-After. Other errors offer an explicit retry.
Concurrent matches in the same organization share server-side per-model pacing;
a busy Mercury queue does not pause other models. Gateway quotas remain authoritative.
Pause finishes current moves and stops new calls; switching tabs pauses the match.
Reset starts over. Match state lives only in the open page.

Settings offers five **Block sequences**: each is a repeatable order of pieces shared
by all models. The boards keep a fixed order: Jev, Perplexity, Bespoke, Clef, Clef-flash,
Mercury, Kev, Decider. **Standings** rank the models by score while the match runs. A row
moves when its model passes another, equal scores share a place, and a line clear or a
change of place shows for a moment. Standings compare scores at each model's current
progress; final standings appear when every board finishes or tops out. Both modes use a
10 × 20 board and seven-bag pieces, line-clear and level points, hard-drop points, combos,
back-to-back clears and all-clear bonuses.

Settings → **Moves** picks how a piece reaches its place:

- **Drop and slide** (the default): each piece turns to its orientation at the top,
  then falls; while it falls it may move left or right, so it can tuck sideways into a gap
  under an overhang. It never turns while falling. The model picks any final position the
  piece can reach this way, a straight drop or a tuck. Each choice tells the model which of
  the two it is and the points it scores now; the board shows the piece turn, drift toward
  its column and tuck.
- **Drop only**: the model picks a rotation and a column, and the piece drops straight down.

Hold, wall kicks, T-spins and a gravity deadline are not part of either mode; **Rules &
scoring** states the exact rules.

## Decision

Some models do not write text at all. A *decision model* takes a **state** — a message, a
ticket, a list of items — and typed **questions**, and returns a typed answer for each with
the probability or confidence behind it. It is built for routing, classification and the
decision points inside an application, where a fast predictable answer matters more than
prose. The catalog has ten:

| | Jev 1.13 | Mercury Decide | Bespoke Nimble v3 | Clef and Clef-flash | Decider 2B and Kev 4B | Perplexity Decider v1 27B | Sage 1.2 | GLiNER-2.5-Decide |
|---|---|---|---|---|---|---|---|---|
| Catalog id | `typesafe/jev-1.13` | `inception/mercury-decide` | `bespokelabs/nimble-v3` | `cloudflare/clef`, `cloudflare/clef-flash` | `routerplus/decider-2b`, `routerplus/kev-4b` | `perplexity/pplx-decider-v1-27b` | `levanto/sage-1.2` | `fastino/gliner-2.5-decide` |
| Made by | TypeSafe | Inception | Bespoke Labs | Cloudflare | RouterPlus, our own models | Perplexity | Levanto | Fastino |
| Question format | System One | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | GLiNER's schema |
| Question kinds | choice, rating, yes/no; tags as one yes/no per tag | the same as Jev | the same as Jev | the same as Jev | the same as Jev | the same as Jev | the same as Jev | yes/no, choice, rating, tags, all as labels |
| Rating scale | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | exactly 5 levels |
| Numbers it returns | probabilities | probabilities | probabilities, at full precision | probabilities, to 4 decimals | probabilities | probabilities, at full precision | probabilities | confidences, not calibrated probabilities |
| When unsure | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best label, with a confidence |
| Price | $0.042 per million input tokens | free while OpenRouter serves only its free variant | $0.04 per million input tokens | $0.24 (Clef) or $0.09 (Clef-flash) per million input tokens | free: our own models, at $0 | $0.04 per million input tokens, the state billed once per question | $0.05 per million input tokens, the state billed once per question | $0.03 per million input tokens |

Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and
Sage are the nine **System One** models: they take the same questions and answer in the same
shape, so wherever this page says what System One takes, it means all nine. Nimble, Clef and
Clef-flash have two limits of their own: a choice needs at least 2 options, and one request
holds at most 64 questions, each tag counted as one. Clef and Clef-flash have a third:
every answer key is 1 to 100 letters, digits, `_`, `.` or `-`. A question id in Decision mode
always fits that rule. A tag can break it: Decision mode asks each tag under the key
`<question id>__<tag id>`, and a long question id with a long tag id can pass 100
characters. Decision mode checks these limits before a run. It counts the questions in order,
and only the ones the column will send: a question that would cross 64 shows
**Not askable on Bespoke Nimble v3** (or on Clef, or on Clef-flash) and is left out of that
column's request. It takes none of the 64, and a later question that still fits is sent. So
the provider never refuses the column whole.
For Decider 2B and Kev 4B, Decision mode checks no limit of their own: probed on
2026-10-02, each answers a one-option choice and a request of 200 questions, so Decision mode
holds their columns to the family's rules alone, as it holds a Jev column.
Perplexity Decider has one limit of its own: one request holds at most 128 questions, each
tag counted as one. It takes a one-option choice, as Jev does. Decision mode counts the
questions in order, as for Nimble, and a question that would cross 128 shows **Not askable on
Perplexity Decider v1 27B**. Perplexity bills the state once for each question it asks, so on
a Perplexity Decider column every question, and every tag of a tags question, costs one
state.

Switch to **Decision** and the composer becomes a decision console: a box for what the models
judge, and up to 20 questions that every column is asked. Each question is a line of text
and an answer type: **Yes / no**, **Pick one** (one of your options), **Rating** (a scale,
lowest level first; a new rating starts with five levels) or **Tags** (which of your tags
apply). Options, tags and levels are rows: a short name, which is what the model answers
with, and what it means. You never write an answer key: Decision mode names each question from
its words, and **Code** shows the names it sends. A yes / no saved by the earlier editor with
descriptions of a yes and a no keeps them, shown under the question; only the System One
models take them, and **Remove descriptions** makes the question plain again.

An example loads by itself. The **Example** menu picks another, or **Start blank**. The
examples are the jobs a decision model is for: checking an answer against its sources, reviewing an
agent's run, catching a prompt injection, triage, bug severity, lead qualification, code
review, moderation, feedback tags and a JSON state. **Run** runs it once. **Run 10 times**
runs every column ten times (see below).

### One form for every model

You write each question once. Each decision model has its own question format, so Decision mode
translates the form into that format for each column. This is what each model receives:

| Question | System One (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sage) | GLiNER-2.5-Decide |
|---|---|---|
| Yes / no | a `noul` question | a classification: the question text is the task, the labels are `yes` and `no` |
| Choice, 2 to 20 options, descriptions optional | a `choice` question: each option, with its description or none | a classification: one label per option, `"name: description"` (or `"name"`), read back as the option's name |
| Rating, exactly 5 levels, low to high | a `score` question with 5 levels | a classification: the labels `"0: <level 0>"` to `"4: <level 4>"`, read back as the level |
| Tags, 1 to 20, descriptions optional | one `noul` question per tag, asked as `<question> Does this apply: <tag>?` (with the tag's description after the tag, when it has one) | a multi-label classification that returns every tag with its confidence; a tag applies at 0.5 or more |

GLiNER reads each label as text, so Decision mode puts an option's or level's description into
its label. The description changes GLiNER's answer, as it would for a person. Decision mode asks
GLiNER at threshold 0 (`cls_threshold: 0`), so GLiNER always names its top label, with its
confidence beside it. At GLiNER's default of 0.5 a task answers `null` when no label reaches
it; Decision mode shows a `null` as Not sure. System One has no
tags question, so Decision mode asks a System One model one yes/no per tag. Each of those is a
question in that column's request, and the column shows the tags together. This holds for
every System One model alike (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B,
Kev 4B, Perplexity Decider, Sage): Decision mode builds a System One body from the catalog's word on the model's
format, never from its name.

### Compare

Put up to four models side by side with **Compare** — up to four of the ten decision
models, or three of them and a chat model, or any mix; a chat, image or video conversation
stops at three. One turn cannot hold all ten decision models: pick the four, or fewer, that
you want to compare. A column's chip swaps one model for another.
With two or more decision models in the comparison, the question builder holds to the
table above:

- A question in a form that one of the chosen decision models cannot take is marked before
  the run, with the models that can take it. Examples: a rating that does not have 5
  levels (System One models only), a choice with one option
  (Jev, Mercury Decide, Decider 2B, Kev 4B and Perplexity Decider only: Nimble, Clef and
  Clef-flash refuse a one-option choice), more than 20 options, a question past the 64 keys
  of one Nimble, Clef or Clef-flash request or the 128 keys of one Perplexity Decider
  request, a tag whose key is longer than Clef's 100 characters.
  Models refused for the same reason share one line. The notices name the models in the
  comparison that take the question, so a Mercury Decide, Nimble, Clef, Decider 2B or
  Perplexity Decider column is named as itself, never Jev.
- The run still sends each column the questions its model can take. A question that a
  column's model cannot take shows **Not askable on** that model in the column, and it is
  left out of that column's request. The rest of the questions still go.
- GLiNER needs a different text for each question, because the question text is its task
  name, and a task name is at most 256 characters. Two questions with the same text, or a
  question longer than 256 characters, are not askable on GLiNER, and the notice says so.

The examples use forms that every decision model takes, the tags example included.

### Run 10 times

**Run 10 times** runs the whole comparison ten times — every column, the same state and
questions — and shows one ledger instead of a card per column. Each column sends up to two
requests at a time, and the ledger fills in as the runs come back:

- For each model and question, the answer the model gave most often, and in how many of the
  runs it gave it (**10/10** when it never changed its mind). A pick-one question lists
  every option the runs chose, with how many runs chose it.
- The averages of the model's numbers: the mean probability of yes, the mean confidence of
  a choice, the mean level of a rating (a marker on the track, with a faint dot for each
  run), and for tags how many runs each tag applied in. The bars and markers move to the
  new averages as each run lands.
- At the head of each column, the mean time per run and the cost of all its runs, with each
  run's time as a dot on a strip and the mean as a line.
- Each row says whether the decision models agree on their most common answers (for a
  rating, mean levels within 0.5 of each other).

Each run is billed like a run: ten runs of four columns are 40 requests. A column that hits
a rate limit stops and says so at its head; the other columns go on. **Stop** stops them
all.

On Decider 2B and Kev 4B, the ten runs send the same request, so RouterPlus
answers the repeats from its cache, which keeps an answer for 600 s. Those runs come back
sooner than the first one, cost $0, and repeat the first run's answer exactly: on these
models the ten runs do not show how much the model's answer varies. The first request
after a quiet period can take 15 to 20 s while RouterPlus starts the model.

Below the columns, a compare view draws each question as one row, with one cell per column,
in the same way for every model:

| Question | Each cell shows |
|---|---|
| Yes / no | The verdict (Yes, No or Not sure) and a bar from 0 to 1. For the System One models the bar is the probability of yes. For GLiNER it is its confidence in the label it chose, labelled **confidence**. For a chat model, its stated answer. |
| Choice | The chosen option (or Not sure) and its probability or confidence. Per-option bars where the model returns them: a System One model's add up to 1, GLiNER returns the chosen label only. |
| Rating | One 0 to 4 track, with one marker per model: a System One model's score and GLiNER's chosen level, with the confidence beside it. |
| Tags | A grid of tag by model: applies ✓, does not apply ✗ or not sure ?, with the score. |

Each row says whether the decision models agree: the same verdict, the same option, a level
within 0.5, or the same set of tags, with every tag answered (a tag one model is not sure
about is nothing to compare, and the row says so). The view shows only numbers that a model returned, each
labelled with what it means. It invents none. **GLiNER's numbers are confidences, not
calibrated probabilities**: a GLiNER 0.9 and a Jev 0.9 do not mean the same thing. A Clef
confidence and a Jev confidence do not mean the same thing either: Cloudflare computes Clef's
in its own way, so compare their probabilities, not their confidences. Each
column still shows its own cost and latency. A Mercury Decide, Decider 2B or Kev 4B column's cost is $0.00 with no
list price or discount beside it: the model is free, so there is nothing to take off.

### Answer cards

The answer cards show each model's own answer. A System One model's (Jev's, Mercury
Decide's, Nimble's, Clef's, Clef-flash's, Decider 2B's, Kev 4B's, Perplexity
Decider's and Sage's) show bars: the chosen option, the probability of each alternative,
and the model's confidence. GLiNER's cards show the label it chose and its confidence, and for
tags each tag's confidence.

Add a chat model as a column and the same questions go to it in words, tags
included where the prompt can say them. When it answers in the JSON it was asked for, its
answers render as the same cards and in the compare view — so you can read the typed answer,
the numbers, the latency and the cost side by side.

**A Decision run is a real API call.** It goes through the gateway's
[`POST /v1/decisions`](/docs/api-decisions) under the playground key, like a chat reply. It
is billed to your organization at the model's price, counts against the playground key's
limits and the organization's, and appears in `GET /v1/generation`. Each answer card has an
`audit` link to the request's ledger rows.
Decision models can carry a discount, which may make a run free: the cost line shows what
you paid and, when a discount applies, the list price with the percent off. To call a
model from your code, use [`POST /v1/decisions`](/docs/api-decisions) in that model's own
question format. The translation above happens in Decision mode only; the API does not
translate.

## Images

Pick an image model — GPT Image, Nano Banana, Seedream and Grok Imagine are in the picker,
marked **image** with the most one image can cost — and the composer describes an image
instead of starting a chat. Under the description are the options that model declares:
**size** or **shape**, **resolution**, **quality**, how many **images** (up to four, where the
model allows more than one), **format** and **background**. A model that lacks an option
does not show it. The estimate beside the send button is the most the render can cost: the
per-image ceiling the gateway reserves. The bill is what the render actually cost, which is
usually less.

**Compare** puts up to three image models side by side. The same description goes to each,
with the options it accepts. A comparison holds image models only: choosing a chat model to
lead starts a new conversation, so text and pictures never share a history.

A render usually takes 10 to 60 seconds, and each tile shows the time elapsed. A render
cannot be stopped once it starts — the provider bills it either way — so there is no Stop
button. It keeps going while you open another chat or start a new one: its chat shows a dot
in the list until the picture arrives, and the picture lands there. **A render also
survives a reload**, or a background tab the browser put to sleep: it runs on our side,
and the page takes it up again when it comes back. A finished render waits there for
15 minutes; after that it is gone, though it was still billed. Each image has a **Download** button.
The sample descriptions include the modality's known hard cases: exact spelling,
counting, hands and a transparent background. A sample fills the description rather than
rendering, because a render costs more than a chat reply.

**Images stay in your browser.** They are kept in its own storage (IndexedDB), filed by
your sign-in, the conversation and the reply, so a reload or a visit tomorrow shows them
again. They are never uploaded. A render is metered exactly like
[`POST /v1/images/generations`](/docs/api-images), under the playground key.

## Video

Pick a video model — Seedance 2.5, MiniMax H3 Max or Wan 3.0, marked **video** with the
most one second can cost — and the composer describes a video. Its options come from the
model: **length**, **resolution**, **shape** and, where the model makes sound, **sound**.
The estimate beside the send button is the most the turn can cost: each column's length at
its model's per-second ceiling. The bill is what the provider charged, which is usually
less.

**Compare** puts up to three video models side by side, like images. The description goes
to each with the options it takes: a length longer than a model makes becomes its longest,
and a resolution it lacks becomes its own default. Each column's card says what it was
asked for.

A video takes from ten seconds to several minutes. Its tile says where the job is —
**Queued at the provider**, then **Being made** — and for how long. You can leave: open
another chat, reload, or close the tab and come back later. The job keeps going at the
provider, and the page follows it again when you return. There is no Stop: the provider
bills a job once it starts. When it is done the video plays in the card, with its cost and
a **Download** button, and it is kept in your browser like an image. A video is metered
exactly like [`POST /v1/videos`](/docs/api-videos), under the playground key.

## How it is billed

Playground usage is metered exactly like an API call. The console mints one key named
**Playground** per organization, tagged **managed** on
[`/console/keys`](https://app.routerplus.com/console/keys) (turn on **Show playground keys** there to
list it; the page hides managed keys by default); every reply is an attempt under that key,
charged to the organization's credits, subject to its spend caps, and visible in the
console and `GET /v1/generation`. BYOK connections bound to that key apply too.

The key's value is never shown to anyone: the console holds it in memory only and calls the
gateway on your behalf. The key is replaced every 24 hours and after a console restart; the
one it replaces is retired. Disabling it from the console just makes the playground mint a
new one on the next message.

## What is stored

Nothing content-bearing on our side. Conversations live in your browser's local storage,
up to 100 of them: use **Export** and **Import** to move them, **Clear all** to delete
them. Images and videos live in your browser's IndexedDB storage, up to 1 GB before the
oldest go first; deleting a chat, or **Clear all**, deletes its files too. Export carries the
conversations, not the files — download the ones you want to keep elsewhere. The
marketplace records only what it records for every request: model, provider, timestamps,
token counts and cost. For a video it also keeps the job — its id, model, length, shape,
status and charge — so the page can follow it; never the prompt or the file. See
[Data policy](/docs/data-policy).

## Limits

- No attachments. A chat takes text; an image or video model takes a text description.
  Image editing, image-to-video and reference images are not supported yet.
- Up to three image renders in flight per organization at once — one per comparison
  column — and up to four images per render. A video is a job at the gateway: your
  balance must cover each column's reservation.
- A conversation is capped at 400 000 characters per request and 200 messages, and a
  system prompt at 32 000 characters; start a new chat when a model reports a context
  overflow.
- The managed key has its own request rate limit (60 per minute) on top of the organization's
  limits. A comparison spends one request per column, and the gateway holds a credit estimate
  for each: with a low balance, one column can be refused while another answers. A Decision
  run counts the same way, and **Run 10 times** sends ten requests per column, two at a
  time per column.
- Columns served by the same provider share that provider's pool. There is room for a
  full comparison — three chat models, or the four columns in Decision mode — and for other people
  working at the same time; a comparison wide
  enough to exhaust it returns `rate_limit` on the later columns, whose card offers
  **Try again**. Mercury Decide's pool is small (20 requests a minute in all, 15 per
  organization, and 5 in flight at once, 3 per organization, because OpenRouter serves only
  its free variant), and OpenRouter caps free
  requests per day across our account: **Run 10 times** with Mercury Decide can meet
  either limit, and its column says **rate limited** when it does. Nimble's pool takes 7 requests
  at once in all, 5 per organization: Bespoke allows our whole account 8, and one is kept
  for our test environment. Clef and Clef-flash share one pool: 200 requests a minute and
  at most 8 at a time in all, 150 and 6 per organization. Decider 2B and Kev 4B
  share one RouterPlus pool, which is large (11,400 requests a minute and 59 in flight at
  once in all, 44 per organization), so your organization's own limits bind first.
  Perplexity Decider's pool takes 540 requests a minute and 8 in flight at once in all,
  405 and 6 per organization: Perplexity allows our organization 10 requests a second.
- Each model carries the mark of the lab that made it, in that lab's own colours, or the
  model's own mark where it has one. A model id the catalog does not list gets a neutral
  chip rather than a guessed logo.

## Errors

Errors show the same class and remediation as the API's [error table](/docs/errors): a
`rate_limit` asks you to wait, `insufficient_quota` points you at credits, `content_policy`
is never rerouted to another provider. Every failed reply carries its request id.
