Let your agent make the model-buying call
Choosing a model is usually a judgment call. Someone picks one that seems good enough, tries it on a few inputs, and commits. Framing it as a procurement decision changes the question from “which model is best” to something an agent can answer with data: which model is the cheapest that clears a quality bar on the actual task.
And the answer does not have to be one model. There is no rule that a stack runs on a single provider. Once the choice is driven by a quality bar instead of a default, each step in a pipeline can run on the cheapest model that clears the bar for that step. Extraction on a budget model, reasoning on a frontier one, classification on whichever open-source model scores highest per dollar. The monolithic AI stack is a habit, not a requirement, and it is usually the expensive habit.
Set the bar
A quality bar is a score a model has to reach to be considered at all. It turns a ranking problem into a filter: models that clear the bar are candidates, models that do not are out, and among the candidates the cheapest wins.
In LLM Prover, the bar is implemented as a rubric score. The task here: pull the action items out of a messy meeting note into clean JSON a downstream system can route. Here is the exact input every model received.
The prompt
Extract the action items from this meeting note as a JSON array, each with fields ‘owner’ and ’task’. Return only the JSON.
Note: Priya to send the revised pricing deck to legal by Thursday. Marcus said he’d loop in the design team once the copy is locked. We still need someone to confirm the Q3 numbers with finance before the board call. Priya volunteered to own that too.
The rubric that scored each answer:
Rubric (what the judge scores)
- Valid JSON array (25%): an array of objects with exactly ‘owner’ and ’task’ keys, no prose, no code fences.
- Correct owners (30%): every item names the right person from the note.
- Correct tasks (30%): every item describes the right task, nothing invented, nothing missed.
- Concise descriptions (15%): each task is one direct sentence.
The bar for this task: 80 out of 100. Good enough to put into a pipeline that routes action items to owners.
Score the field
The agent ran all four candidate models against the same note with the same rubric.
| Model | Quality | Cost | Clears bar (80) |
|---|---|---|---|
| GPT-4.1 | 100 | $0.00086 | yes |
| DeepSeek V4 Flash | 100 | $0.00008 | yes |
| GPT-4o Mini | 75 | $0.00007 | no |
| Claude Haiku 4.5 | 75 | $0.00065 | no |
Two models cleared the bar. The two that missed both lost the same point.
The two that scored 75 extracted every owner and task correctly. They lost the JSON-structure criterion for the same reason: both wrapped the output in markdown code fences, which a downstream parser would choke on. Correct content, wrong envelope. For a pipeline that reads the JSON directly, that is a fail, and the rubric caught it.
The recommendation writes itself
Two models cleared the bar at a perfect score: GPT-4.1 and DeepSeek V4 Flash. One costs $0.00086 per call, the other $0.00008. Same score, 11x price difference.
The agent does not need to deliberate. The bar eliminated two models, the cost column ranked the survivors, and the recommendation is the cheapest model that cleared the bar: DeepSeek V4 Flash. GPT-4.1 is an excellent model and scored identically, but on this task, at this bar, it is 11x more expensive for the same result.
The agent is not picking the best model. It is removing every model that fails the bar and recommending the cheapest one left. That is a procurement decision, and it is the kind of call that quietly saves the most money, because the expensive default often scores no better than the budget option on the task you actually run.
Put a quality bar on your model choice
Score your candidates against your task and let the numbers decide. Pro and above.
When the recommendation changes
This result holds for this task, this rubric, and this bar. Raise the bar to require perfect JSON structure and the field is unchanged. Change the task to something harder and the two budget models may separate from the frontier ones. A model the provider updates next month may drop below the bar it clears today.
That is why the recommendation is a measurement, not a verdict. Run it again when the task changes, the bar changes, or a model ships a new version. For catching that last case automatically, see how a scheduled benchmark flags a model that regresses. For the full model-selection methodology behind the bar, see How to find the right LLM for your use case.
Picking the model is only the first saving
Choosing the cheapest model that clears the bar is one lever. The prompt is another. The same MCP tools that score a field of models can run an optimization loop that tunes a prompt against the rubric, iteration by iteration, keeping the version that scores higher. A tighter prompt often lets a cheaper model clear a bar it missed before, which pushes the whole decision further down the cost curve.
So the two levers compound. First the agent finds the cheapest model that passes. Then it optimizes the prompt so an even cheaper model passes, or so the passing model runs on fewer tokens. Both run on the same tools, both driven by the same quality bar, both the agent’s work rather than yours. The optimization loop shows that second lever in action.
What’s next
MCP Recipes Guide
Hand your AI agent a ready-to-paste playbook that runs LLM Prover tools in order, polls for results, and reports back in plain language.
Rubric Guide
How to write criteria that produce reliable, consistent judge scores, and how to fix them when they don't.
Comparisons Guide
Fire a prompt at multiple models simultaneously and compare cost, latency, and output side by side.