Playground
Chat, decision, images and video with any catalog model; what it costs and what is stored.
https://app.routerplus.com/playground is an interface over the same catalog, routing, prices and ledger as the API: chat, typed decisions, images and video. Sign in, pick a model, and each answer arrives with the provider that served it, the token counts, and what it cost.
What it does
- Any catalog model. The picker searches the catalog;
⌘K/Ctrl+Kopens it. It lists the models of the mode you are in first (chat models in Chat, decision models in Decision). Chat starts with GPT-6 Astra, Claude Opus 5.5, GLM 5.3 Flash, DeepSeek V4.1 Flash and Hy4 Preview; then the other open-weight models; then GPT-6.1 Sol, GPT-6 Luna and Claude Sonnet 5.5; then the rest, each group newest first. Decision mode lists Jev 1.13, Bespoke Nimble v3, Clef, Clef-flash, Mercury Decide, Kev 4B, GLiNER-2.5-Decide, Decider 2B and Sage 1.2, in that order. The search matches a model's name, id, lab, the providers that run it, and its kind: typedecisionfor every decision model,imageorvideofor those,freefor the free ones. Each row carries the lab's mark, its price — per million tokens, or for an image or video model the provider's own price and the most one image or second can cost — and a chat model's context length. Typing an id that is not listed offers it verbatim, which is how a dated snapshot of a catalog model is pinned. - Compare up to three chat models side by side (Decision mode compares up to four). Compare adds a column. One prompt goes to every column at once, each streams into its own card, and each shows its own provider, tokens, latency and cost. Every column keeps its own thread: a follow-up sees that model's replies only, never another's. Removing a column leaves the rest untouched.
- Sample prompts under the composer send a real prompt in one click. The first is always the strawberry test (how many times "r" appears in "strawberry"), asked plainly. Each reply to it says passed, failed with the count the model gave, or that it gave no clear count. The others are drawn at random: a reasoning trap, a code task, structured JSON, a table, a six-word story and a few more. They make way as soon as the conversation starts.
- Streaming replies with a stop button, regenerate, a collapsible "thinking" section for reasoning models, and safe rendering of code blocks, lists, links and tables. A table is drawn only when the reply holds a real markdown table: a header row, then a divider row such as
|---|---|with the same number of cells. Any other text with a pipe in it reads as it is. A wide table scrolls sideways inside its card. - Code and Parameters sit in the composer, to the left of Send. Code shows the request the playground is about to make, as curl, Python or Node, so you can paste it into your own project. Parameters holds a system prompt, temperature, top_p and max output tokens, per conversation.
- Cost on every reply: the same
usage.costthe API returns, plus time to first byte and anauditlink to the request's ledger rows. - Organizations: if you belong to several, the sidebar selector switches which one the playground bills.
Tetris Royale
Tetris Royale compares the eight current System One decision models on the same block sequence. Open it from the Playground mode bar; the arrow button at the top left goes back. It keeps the selected organization and uses a separate server-held managed key named Tetris Royale; ordinary Playground keys retain their existing limit. The server validates the board and match counters, then constructs the decision prompt and legal choices itself. The game relay does not accept custom prompts or questions. Every model request follows the organization's normal credits, limits, billing and provider routing. No API key reaches the browser.
A match defaults to 200 pieces; Settings offers 20 up to a thousand. Each model plays independently: a slow or rate-limited model does not hold up the others. Transient 429s cool down only that model and retry up to three times, respecting Retry-After. Other errors offer an explicit retry. Concurrent matches in the same organization share server-side per-model pacing; a busy Mercury queue does not pause other models. Gateway quotas remain authoritative. Pause finishes current moves and stops new calls; switching tabs pauses the match. Reset starts over. Match state lives only in the open page.
Settings offers five Block sequences: each is a repeatable order of pieces shared by all models. The boards keep a fixed order: Jev, Perplexity, Bespoke, Clef, Clef-flash, Mercury, Kev, Decider. Standings rank the models by score while the match runs. A row moves when its model passes another, equal scores share a place, and a line clear or a change of place shows for a moment. Standings compare scores at each model's current progress; final standings appear when every board finishes or tops out. Both modes use a 10 × 20 board and seven-bag pieces, line-clear and level points, hard-drop points, combos, back-to-back clears and all-clear bonuses.
Settings → Moves picks how a piece reaches its place:
- Drop and slide (the default): each piece turns to its orientation at the top, then falls; while it falls it may move left or right, so it can tuck sideways into a gap under an overhang. It never turns while falling. The model picks any final position the piece can reach this way, a straight drop or a tuck. Each choice tells the model which of the two it is and the points it scores now; the board shows the piece turn, drift toward its column and tuck.
- Drop only: the model picks a rotation and a column, and the piece drops straight down.
Hold, wall kicks, T-spins and a gravity deadline are not part of either mode; Rules & scoring states the exact rules.
Decision
Some models do not write text at all. A decision model takes a state — a message, a ticket, a list of items — and typed questions, and returns a typed answer for each with the probability or confidence behind it. It is built for routing, classification and the decision points inside an application, where a fast predictable answer matters more than prose. The catalog has ten:
| Jev 1.13 | Mercury Decide | Bespoke Nimble v3 | Clef and Clef-flash | Decider 2B and Kev 4B | Perplexity Decider v1 27B | Sage 1.2 | GLiNER-2.5-Decide | |
|---|---|---|---|---|---|---|---|---|
| Catalog id | typesafe/jev-1.13 | inception/mercury-decide | bespokelabs/nimble-v3 | cloudflare/clef, cloudflare/clef-flash | routerplus/decider-2b, routerplus/kev-4b | perplexity/pplx-decider-v1-27b | levanto/sage-1.2 | fastino/gliner-2.5-decide |
| Made by | TypeSafe | Inception | Bespoke Labs | Cloudflare | RouterPlus, our own models | Perplexity | Levanto | Fastino |
| Question format | System One | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | System One, the same as Jev | GLiNER's schema |
| Question kinds | choice, rating, yes/no; tags as one yes/no per tag | the same as Jev | the same as Jev | the same as Jev | the same as Jev | the same as Jev | the same as Jev | yes/no, choice, rating, tags, all as labels |
| Rating scale | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | 2 to 10 levels | exactly 5 levels |
| Numbers it returns | probabilities | probabilities | probabilities, at full precision | probabilities, to 4 decimals | probabilities | probabilities, at full precision | probabilities | confidences, not calibrated probabilities |
| When unsure | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best answer, with probabilities | gives its best label, with a confidence |
| Price | $0.042 per million input tokens | free while OpenRouter serves only its free variant | $0.04 per million input tokens | $0.24 (Clef) or $0.09 (Clef-flash) per million input tokens | free: our own models, at $0 | $0.04 per million input tokens, the state billed once per question | $0.05 per million input tokens, the state billed once per question | $0.03 per million input tokens |
Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider and Sage are the nine System One models: they take the same questions and answer in the same shape, so wherever this page says what System One takes, it means all nine. Nimble, Clef and Clef-flash have two limits of their own: a choice needs at least 2 options, and one request holds at most 64 questions, each tag counted as one. Clef and Clef-flash have a third: every answer key is 1 to 100 letters, digits, _, . or -. A question id in Decision mode always fits that rule. A tag can break it: Decision mode asks each tag under the key <question id>__<tag id>, and a long question id with a long tag id can pass 100 characters. Decision mode checks these limits before a run. It counts the questions in order, and only the ones the column will send: a question that would cross 64 shows Not askable on Bespoke Nimble v3 (or on Clef, or on Clef-flash) and is left out of that column's request. It takes none of the 64, and a later question that still fits is sent. So the provider never refuses the column whole. For Decider 2B and Kev 4B, Decision mode checks no limit of their own: probed on 2026-10-02, each answers a one-option choice and a request of 200 questions, so Decision mode holds their columns to the family's rules alone, as it holds a Jev column. Perplexity Decider has one limit of its own: one request holds at most 128 questions, each tag counted as one. It takes a one-option choice, as Jev does. Decision mode counts the questions in order, as for Nimble, and a question that would cross 128 shows Not askable on Perplexity Decider v1 27B. Perplexity bills the state once for each question it asks, so on a Perplexity Decider column every question, and every tag of a tags question, costs one state.
Switch to Decision and the composer becomes a decision console: a box for what the models judge, and up to 20 questions that every column is asked. Each question is a line of text and an answer type: Yes / no, Pick one (one of your options), Rating (a scale, lowest level first; a new rating starts with five levels) or Tags (which of your tags apply). Options, tags and levels are rows: a short name, which is what the model answers with, and what it means. You never write an answer key: Decision mode names each question from its words, and Code shows the names it sends. A yes / no saved by the earlier editor with descriptions of a yes and a no keeps them, shown under the question; only the System One models take them, and Remove descriptions makes the question plain again.
An example loads by itself. The Example menu picks another, or Start blank. The examples are the jobs a decision model is for: checking an answer against its sources, reviewing an agent's run, catching a prompt injection, triage, bug severity, lead qualification, code review, moderation, feedback tags and a JSON state. Run runs it once. Run 10 times runs every column ten times (see below).
One form for every model
You write each question once. Each decision model has its own question format, so Decision mode translates the form into that format for each column. This is what each model receives:
| Question | System One (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sage) | GLiNER-2.5-Decide |
|---|---|---|
| Yes / no | a noul question | a classification: the question text is the task, the labels are yes and no |
| Choice, 2 to 20 options, descriptions optional | a choice question: each option, with its description or none | a classification: one label per option, "name: description" (or "name"), read back as the option's name |
| Rating, exactly 5 levels, low to high | a score question with 5 levels | a classification: the labels "0: <level 0>" to "4: <level 4>", read back as the level |
| Tags, 1 to 20, descriptions optional | one noul question per tag, asked as <question> Does this apply: <tag>? (with the tag's description after the tag, when it has one) | a multi-label classification that returns every tag with its confidence; a tag applies at 0.5 or more |
GLiNER reads each label as text, so Decision mode puts an option's or level's description into its label. The description changes GLiNER's answer, as it would for a person. Decision mode asks GLiNER at threshold 0 (cls_threshold: 0), so GLiNER always names its top label, with its confidence beside it. At GLiNER's default of 0.5 a task answers null when no label reaches it; Decision mode shows a null as Not sure. System One has no tags question, so Decision mode asks a System One model one yes/no per tag. Each of those is a question in that column's request, and the column shows the tags together. This holds for every System One model alike (Jev, Mercury Decide, Nimble, Clef, Clef-flash, Decider 2B, Kev 4B, Perplexity Decider, Sage): Decision mode builds a System One body from the catalog's word on the model's format, never from its name.
Compare
Put up to four models side by side with Compare — up to four of the ten decision models, or three of them and a chat model, or any mix; a chat, image or video conversation stops at three. One turn cannot hold all ten decision models: pick the four, or fewer, that you want to compare. A column's chip swaps one model for another. With two or more decision models in the comparison, the question builder holds to the table above:
- A question in a form that one of the chosen decision models cannot take is marked before the run, with the models that can take it. Examples: a rating that does not have 5 levels (System One models only), a choice with one option (Jev, Mercury Decide, Decider 2B, Kev 4B and Perplexity Decider only: Nimble, Clef and Clef-flash refuse a one-option choice), more than 20 options, a question past the 64 keys of one Nimble, Clef or Clef-flash request or the 128 keys of one Perplexity Decider request, a tag whose key is longer than Clef's 100 characters. Models refused for the same reason share one line. The notices name the models in the comparison that take the question, so a Mercury Decide, Nimble, Clef, Decider 2B or Perplexity Decider column is named as itself, never Jev.
- The run still sends each column the questions its model can take. A question that a column's model cannot take shows Not askable on that model in the column, and it is left out of that column's request. The rest of the questions still go.
- GLiNER needs a different text for each question, because the question text is its task name, and a task name is at most 256 characters. Two questions with the same text, or a question longer than 256 characters, are not askable on GLiNER, and the notice says so.
The examples use forms that every decision model takes, the tags example included.
Run 10 times
Run 10 times runs the whole comparison ten times — every column, the same state and questions — and shows one ledger instead of a card per column. Each column sends up to two requests at a time, and the ledger fills in as the runs come back:
- For each model and question, the answer the model gave most often, and in how many of the runs it gave it (10/10 when it never changed its mind). A pick-one question lists every option the runs chose, with how many runs chose it.
- The averages of the model's numbers: the mean probability of yes, the mean confidence of a choice, the mean level of a rating (a marker on the track, with a faint dot for each run), and for tags how many runs each tag applied in. The bars and markers move to the new averages as each run lands.
- At the head of each column, the mean time per run and the cost of all its runs, with each run's time as a dot on a strip and the mean as a line.
- Each row says whether the decision models agree on their most common answers (for a rating, mean levels within 0.5 of each other).
Each run is billed like a run: ten runs of four columns are 40 requests. A column that hits a rate limit stops and says so at its head; the other columns go on. Stop stops them all.
On Decider 2B and Kev 4B, the ten runs send the same request, so RouterPlus answers the repeats from its cache, which keeps an answer for 600 s. Those runs come back sooner than the first one, cost $0, and repeat the first run's answer exactly: on these models the ten runs do not show how much the model's answer varies. The first request after a quiet period can take 15 to 20 s while RouterPlus starts the model.
Below the columns, a compare view draws each question as one row, with one cell per column, in the same way for every model:
| Question | Each cell shows |
|---|---|
| Yes / no | The verdict (Yes, No or Not sure) and a bar from 0 to 1. For the System One models the bar is the probability of yes. For GLiNER it is its confidence in the label it chose, labelled confidence. For a chat model, its stated answer. |
| Choice | The chosen option (or Not sure) and its probability or confidence. Per-option bars where the model returns them: a System One model's add up to 1, GLiNER returns the chosen label only. |
| Rating | One 0 to 4 track, with one marker per model: a System One model's score and GLiNER's chosen level, with the confidence beside it. |
| Tags | A grid of tag by model: applies ✓, does not apply ✗ or not sure ?, with the score. |
Each row says whether the decision models agree: the same verdict, the same option, a level within 0.5, or the same set of tags, with every tag answered (a tag one model is not sure about is nothing to compare, and the row says so). The view shows only numbers that a model returned, each labelled with what it means. It invents none. GLiNER's numbers are confidences, not calibrated probabilities: a GLiNER 0.9 and a Jev 0.9 do not mean the same thing. A Clef confidence and a Jev confidence do not mean the same thing either: Cloudflare computes Clef's in its own way, so compare their probabilities, not their confidences. Each column still shows its own cost and latency. A Mercury Decide, Decider 2B or Kev 4B column's cost is $0.00 with no list price or discount beside it: the model is free, so there is nothing to take off.
Answer cards
The answer cards show each model's own answer. A System One model's (Jev's, Mercury Decide's, Nimble's, Clef's, Clef-flash's, Decider 2B's, Kev 4B's, Perplexity Decider's and Sage's) show bars: the chosen option, the probability of each alternative, and the model's confidence. GLiNER's cards show the label it chose and its confidence, and for tags each tag's confidence.
Add a chat model as a column and the same questions go to it in words, tags included where the prompt can say them. When it answers in the JSON it was asked for, its answers render as the same cards and in the compare view — so you can read the typed answer, the numbers, the latency and the cost side by side.
A Decision run is a real API call. It goes through the gateway's POST /v1/decisions under the playground key, like a chat reply. It is billed to your organization at the model's price, counts against the playground key's limits and the organization's, and appears in GET /v1/generation. Each answer card has an audit link to the request's ledger rows. Decision models can carry a discount, which may make a run free: the cost line shows what you paid and, when a discount applies, the list price with the percent off. To call a model from your code, use POST /v1/decisions in that model's own question format. The translation above happens in Decision mode only; the API does not translate.
Images
Pick an image model — GPT Image, Nano Banana, Seedream and Grok Imagine are in the picker, marked image with the most one image can cost — and the composer describes an image instead of starting a chat. Under the description are the options that model declares: size or shape, resolution, quality, how many images (up to four, where the model allows more than one), format and background. A model that lacks an option does not show it. The estimate beside the send button is the most the render can cost: the per-image ceiling the gateway reserves. The bill is what the render actually cost, which is usually less.
Compare puts up to three image models side by side. The same description goes to each, with the options it accepts. A comparison holds image models only: choosing a chat model to lead starts a new conversation, so text and pictures never share a history.
A render usually takes 10 to 60 seconds, and each tile shows the time elapsed. A render cannot be stopped once it starts — the provider bills it either way — so there is no Stop button. It keeps going while you open another chat or start a new one: its chat shows a dot in the list until the picture arrives, and the picture lands there. A render also survives a reload, or a background tab the browser put to sleep: it runs on our side, and the page takes it up again when it comes back. A finished render waits there for 15 minutes; after that it is gone, though it was still billed. Each image has a Download button. The sample descriptions include the modality's known hard cases: exact spelling, counting, hands and a transparent background. A sample fills the description rather than rendering, because a render costs more than a chat reply.
Images stay in your browser. They are kept in its own storage (IndexedDB), filed by your sign-in, the conversation and the reply, so a reload or a visit tomorrow shows them again. They are never uploaded. A render is metered exactly like POST /v1/images/generations, under the playground key.
Video
Pick a video model — Seedance 2.5, MiniMax H3 Max or Wan 3.0, marked video with the most one second can cost — and the composer describes a video. Its options come from the model: length, resolution, shape and, where the model makes sound, sound. The estimate beside the send button is the most the turn can cost: each column's length at its model's per-second ceiling. The bill is what the provider charged, which is usually less.
Compare puts up to three video models side by side, like images. The description goes to each with the options it takes: a length longer than a model makes becomes its longest, and a resolution it lacks becomes its own default. Each column's card says what it was asked for.
A video takes from ten seconds to several minutes. Its tile says where the job is — Queued at the provider, then Being made — and for how long. You can leave: open another chat, reload, or close the tab and come back later. The job keeps going at the provider, and the page follows it again when you return. There is no Stop: the provider bills a job once it starts. When it is done the video plays in the card, with its cost and a Download button, and it is kept in your browser like an image. A video is metered exactly like POST /v1/videos, under the playground key.
How it is billed
Playground usage is metered exactly like an API call. The console mints one key named Playground per organization, tagged managed on /console/keys (turn on Show playground keys there to list it; the page hides managed keys by default); every reply is an attempt under that key, charged to the organization's credits, subject to its spend caps, and visible in the console and GET /v1/generation. BYOK connections bound to that key apply too.
The key's value is never shown to anyone: the console holds it in memory only and calls the gateway on your behalf. The key is replaced every 24 hours and after a console restart; the one it replaces is retired. Disabling it from the console just makes the playground mint a new one on the next message.
What is stored
Nothing content-bearing on our side. Conversations live in your browser's local storage, up to 100 of them: use Export and Import to move them, Clear all to delete them. Images and videos live in your browser's IndexedDB storage, up to 1 GB before the oldest go first; deleting a chat, or Clear all, deletes its files too. Export carries the conversations, not the files — download the ones you want to keep elsewhere. The marketplace records only what it records for every request: model, provider, timestamps, token counts and cost. For a video it also keeps the job — its id, model, length, shape, status and charge — so the page can follow it; never the prompt or the file. See Data policy.
Limits
- No attachments. A chat takes text; an image or video model takes a text description. Image editing, image-to-video and reference images are not supported yet.
- Up to three image renders in flight per organization at once — one per comparison column — and up to four images per render. A video is a job at the gateway: your balance must cover each column's reservation.
- A conversation is capped at 400 000 characters per request and 200 messages, and a system prompt at 32 000 characters; start a new chat when a model reports a context overflow.
- The managed key has its own request rate limit (60 per minute) on top of the organization's limits. A comparison spends one request per column, and the gateway holds a credit estimate for each: with a low balance, one column can be refused while another answers. A Decision run counts the same way, and Run 10 times sends ten requests per column, two at a time per column.
- Columns served by the same provider share that provider's pool. There is room for a full comparison — three chat models, or the four columns in Decision mode — and for other people working at the same time; a comparison wide enough to exhaust it returns
rate_limiton the later columns, whose card offers Try again. Mercury Decide's pool is small (20 requests a minute in all, 15 per organization, and 5 in flight at once, 3 per organization, because OpenRouter serves only its free variant), and OpenRouter caps free requests per day across our account: Run 10 times with Mercury Decide can meet either limit, and its column says rate limited when it does. Nimble's pool takes 7 requests at once in all, 5 per organization: Bespoke allows our whole account 8, and one is kept for our test environment. Clef and Clef-flash share one pool: 200 requests a minute and at most 8 at a time in all, 150 and 6 per organization. Decider 2B and Kev 4B share one RouterPlus pool, which is large (11,400 requests a minute and 59 in flight at once in all, 44 per organization), so your organization's own limits bind first. Perplexity Decider's pool takes 540 requests a minute and 8 in flight at once in all, 405 and 6 per organization: Perplexity allows our organization 10 requests a second. - Each model carries the mark of the lab that made it, in that lab's own colours, or the model's own mark where it has one. A model id the catalog does not list gets a neutral chip rather than a guessed logo.
Errors
Errors show the same class and remediation as the API's error table: a rate_limit asks you to wait, insufficient_quota points you at credits, content_policy is never rerouted to another provider. Every failed reply carries its request id.
Markdown source for agents: /docs/playground.md · index at /llms.txt