Optimize
Find a cheaper model or mix that scores as well as yours on your Braintrust evals, and deploy it.
Optimize finds a cheaper model or mix that scores as well as yours on your own Braintrust evals, and deploys it as an endpoint: one model id you call like any other.
Before you start
- 01Connect Braintrust on Integrations. Paste a Braintrust API key. The key is checked with Braintrust and stored encrypted, and the page shows only its last four characters. Owners and admins can connect, replace or disconnect it; viewers can only look. (Langfuse is on the same page as Request access: it is set up with the team, not self-serve.)
- 02In Braintrust, have one of these to search on:
| Source | What the search uses |
|---|---|
| An experiment that ran on a dataset | The dataset is the test cases. The experiment's metadata.model is the baseline. Its prompt comes from metadata.prompt_id or metadata.prompt_slug, or the project's only prompt. |
| A dataset | The rows are the test cases. You pick the prompt, and a baseline model if you want one. Without one, the strongest candidate becomes the baseline. |
| Production logs | Each logged call is a test case, with the production reply as the expected turn. A candidate is scored on whether its reply is as good as production's. The baseline is the model most of the logged calls used, unless you pick one. |
The baseline must be a catalog model or one of your own deployed endpoints (tm/...). A search takes up to 2,000 cases.
For an experiment or a dataset, these parts of the rows matter:
| Braintrust | Used as |
|---|---|
A row's input.messages | One test of the next assistant turn. |
| The project's scorers | Quality. A case's score is the mean of its scorers. |
Row tags | The slices in a result's detail view. |
Row metadata.conversation_id and metadata.goal | Keeps the turns of one conversation together. The goal is used for simulated customers. |
Code scorers run in Braintrust. Prompt scorers (LLM graders) run through the gateway and are billed to your credits.
Start a search
Click New search at the bottom of the searches list. Pick the source, name the search, and set the search budget. The budget is the most the search spends on model calls, from $5 to $100. Model calls are billed to your credits as the search runs, under a managed key made for the search (turn on Show playground keys on /console/keys to see it, with a managed tag). Viewers cannot start or change a search.
The search runs in stages. The budget pays for them in this order:
- 01Shortlist. Public benchmarks pick the open models to try, as the Pareto frontier of benchmark score against price. Each model is tried on up to three of its hosts: the cheapest, the one at the highest precision, and the model maker's own.
- 02Pilot. Every candidate answers a sample of the conversations.
- 03Dev. The leaders answer all the dev conversations.
- 04Mixes. Cascades, routers and votes are built from the saved answers; a search agent proposes mixes to try. These cost nothing extra.
- 05Held-out. The finalists, including a fusion of the best models, answer conversations the search never saw. The winner is picked on these.
- 06Customers. The finalists talk to a simulated customer who has the goal of a held-out conversation.
When the search reaches its budget it stops and keeps its results. Add $25 continues it from where it stopped. Stop ends it and keeps its results.
If a Braintrust read fails while the search loads its cases, the worker tries again 3 more times — after 2 s, 10 s and 30 s — before it stops the search. A stopped search says why; fix the cause (reconnect the key, say) and add budget to continue.
Read the results
The chart plots each candidate's cost against its score, with the best-value frontier as a line. The table below it:
| Column | Meaning |
|---|---|
| Model | The model, or tm-mix-N for a mix. Open a row to see what the mix does. |
| Provider | The host that served the model, with its precision when the host states it. |
| Score | The mean scorer score, in percent. |
| $ / 1k req | The measured cost of 1,000 requests like your cases. |
| p50 | The median time to a full answer. |
| TTFT | The median time to the first token. |
| Out tok | Output tokens per answer. |
| Status | Which stage scored the row: pilot, dev or held-out. |
The ★ row is the winner: the cheapest finalist that costs less than the baseline, is not worse than the baseline on a paired test over the same held-out cases, and scores within 2 points of it. Ties on cost go to the lower median latency. Rows marked dev or pilot were scored on fewer conversations and cannot win. If no row qualifies, there is no star.
Open a row to see the cases where it did worse than the baseline, side by side, and its score per tag.
Deploy the result
Deploy makes a result callable as tm/<name>-v1. The name is lowercase letters, digits and dashes, 40 at most:
curl https://api.routerplus.com/v1/chat/completions \
-H "Authorization: Bearer $TM_KEY" \
-H "content-type: application/json" \
-d '{"model": "tm/billing-agent-v1", "messages": [{"role": "user", "content": "Hi"}]}'- Only keys of your organization can call it.
- Each model call it makes is billed like a direct call to that model. The response's
usage.costis their sum. - Send it on
POST /v1/chat/completions. With"stream": trueit returns the finished answer as one chunk, then usage. - Deploying again gives a new version (
-v2). Earlier versions keep working.
Endpoints lists everything your organization deployed, with the search it came from. Try it sends one held-out case of that search to the endpoint and bills your credits like any call.
What is stored
A search stores its cases and the answers candidates gave, for your organization only, until you delete the search. Deleting a search deletes them. A deployed model id keeps working after its search is deleted.
Markdown source for agents: /docs/model-search.md · index at /llms.txt