# Optimize

[Optimize](https://app.routerplus.com/model-search) finds a cheaper model or mix that scores as well as yours on your own Braintrust evals, and deploys it as an endpoint: one model id you call like any other.

## Before you start

1. Connect Braintrust on [Integrations](https://app.routerplus.com/console/integrations). Paste a Braintrust API key. The key is checked with Braintrust and stored encrypted, and the page shows only its last four characters. Owners and admins can connect, replace or disconnect it; viewers can only look. (Langfuse is on the same page as **Request access**: it is set up with the team, not self-serve.)
2. In Braintrust, have one of these to search on:

| Source | What the search uses |
|---|---|
| An **experiment** that ran on a dataset | The dataset is the test cases. The experiment's `metadata.model` is the baseline. Its prompt comes from `metadata.prompt_id` or `metadata.prompt_slug`, or the project's only prompt. |
| A **dataset** | The rows are the test cases. You pick the prompt, and a baseline model if you want one. Without one, the strongest candidate becomes the baseline. |
| **Production logs** | Each logged call is a test case, with the production reply as the expected turn. A candidate is scored on whether its reply is as good as production's. The baseline is the model most of the logged calls used, unless you pick one. |

The baseline must be a catalog model or one of your own deployed endpoints (`tm/...`). A search takes up to 2,000 cases.

For an experiment or a dataset, these parts of the rows matter:

| Braintrust | Used as |
|---|---|
| A row's `input.messages` | One test of the next assistant turn. |
| The project's scorers | Quality. A case's score is the mean of its scorers. |
| Row `tags` | The slices in a result's detail view. |
| Row `metadata.conversation_id` and `metadata.goal` | Keeps the turns of one conversation together. The goal is used for simulated customers. |

Code scorers run in Braintrust. Prompt scorers (LLM graders) run through the gateway and are billed to your credits.

## Start a search

Click **New search** at the bottom of the searches list. Pick the source, name the search, and set the **search budget**. The budget is the most the search spends on model calls, from $5 to $100. Model calls are billed to your credits as the search runs, under a managed key made for the search (turn on **Show playground keys** on `/console/keys` to see it, with a **managed** tag). Viewers cannot start or change a search.

The search runs in stages. The budget pays for them in this order:

1. **Shortlist.** Public benchmarks pick the open models to try, as the Pareto frontier of benchmark score against price. Each model is tried on up to three of its hosts: the cheapest, the one at the highest precision, and the model maker's own.
2. **Pilot.** Every candidate answers a sample of the conversations.
3. **Dev.** The leaders answer all the dev conversations.
4. **Mixes.** Cascades, routers and votes are built from the saved answers; a search agent proposes mixes to try. These cost nothing extra.
5. **Held-out.** The finalists, including a fusion of the best models, answer conversations the search never saw. The winner is picked on these.
6. **Customers.** The finalists talk to a simulated customer who has the goal of a held-out conversation.

When the search reaches its budget it stops and keeps its results. **Add $25** continues it from where it stopped. **Stop** ends it and keeps its results.

If a Braintrust read fails while the search loads its cases, the worker tries again 3 more times — after 2 s, 10 s and 30 s — before it stops the search. A stopped search says why; fix the cause (reconnect the key, say) and add budget to continue.

## Read the results

The chart plots each candidate's cost against its score, with the best-value frontier as a line. The table below it:

| Column | Meaning |
|---|---|
| Model | The model, or `tm-mix-N` for a mix. Open a row to see what the mix does. |
| Provider | The host that served the model, with its precision when the host states it. |
| Score | The mean scorer score, in percent. |
| $ / 1k req | The measured cost of 1,000 requests like your cases. |
| p50 | The median time to a full answer. |
| TTFT | The median time to the first token. |
| Out tok | Output tokens per answer. |
| Status | Which stage scored the row: `pilot`, `dev` or held-out. |

The ★ row is the winner: the cheapest finalist that costs less than the baseline, is not worse than the baseline on a paired test over the same held-out cases, and scores within 2 points of it. Ties on cost go to the lower median latency. Rows marked `dev` or `pilot` were scored on fewer conversations and cannot win. If no row qualifies, there is no star.

Open a row to see the cases where it did worse than the baseline, side by side, and its score per tag.

## Deploy the result

**Deploy** makes a result callable as `tm/<name>-v1`. The name is lowercase letters, digits and dashes, 40 at most:

```bash
curl https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "tm/billing-agent-v1", "messages": [{"role": "user", "content": "Hi"}]}'
```

- Only keys of your organization can call it.
- Each model call it makes is billed like a direct call to that model. The response's `usage.cost` is their sum.
- Send it on `POST /v1/chat/completions`. With `"stream": true` it returns the finished answer as one chunk, then usage.
- Deploying again gives a new version (`-v2`). Earlier versions keep working.

[Endpoints](https://app.routerplus.com/endpoints) lists everything your organization deployed, with the search it came from. **Try it** sends one held-out case of that search to the endpoint and bills your credits like any call.

## What is stored

A search stores its cases and the answers candidates gave, for your organization only, until you delete the search. Deleting a search deletes them. A deployed model id keeps working after its search is deleted.
