Your LLM stack costs what it costs because you haven't measured it yet.
The spend is recoverable. Most of it is not where teams look first. Three of the seven models in the benchmarks below did not exist twelve months ago, and model pricing has shifted significantly in that time. A stack that was well-optimised then may not be now. The path forward is measurement, and the first thing measurement reveals is that the model is rarely the whole story.
The compounding effect
Pipeline improvements do not just reduce cost on the current model. They change which models are eligible.
A model that cannot clear your quality threshold with a vague prompt and a loose system prompt may clear it once both are tightened, at a fraction of the cost. That is not a one-time saving. Each pass lowers the cost floor and opens the next optimisation.
The improvement loop has five steps:
- Measure: benchmark quality, cost, and latency on your actual task
- Tighten inputs: prompt, system prompt, inference parameters
- Qualify cheaper models: re-run the benchmark with the tightened inputs
- Right-size the architecture: match each task to its minimal model
- Monitor for drift: catch regressions before they reach production
The rest of this post is the evidence for each step.
Input precision is a cost lever
Prompt and system prompt are the same lever. Precise inputs produce shorter, more constrained outputs: fewer tokens, lower cost, and better scores. The switcher below shows all four combinations.
In the bottom-right cell, GPT-4.1 Nano and Claude Opus 5 both scored 100. Nano at $0.13/mo, Claude at $10.23/mo. Same score, 79x cost difference.
Claude Opus 5 -- quality score by input precision
Claude Opus 5 on the same task: vague inputs at $0.01884/call, precise inputs at $0.00341/call. A 5.5x cost reduction. Scores went from 0 to 100.
The 5.5x cost reduction is not a system prompt finding specifically. Precise inputs produced shorter outputs. That is the mechanism. Vagueness is not free.
Parameters are deliberate levers, not defaults
Most teams leave inference parameters at provider defaults. That means the benchmark does not reflect production, and the cost is not being controlled.
Three parameters worth setting deliberately:
- Max tokens. Caps output length directly. Set it to match your production cap and the benchmark predicts production cost accurately. A task with a predictable output length does not need 1000 tokens of headroom.
- Prompt caching. If your benchmark uses a fixed system prompt across many runs, caching reduces repeated prefix costs on supported providers. Set it once and it applies to every run.
- Temperature and seed. Low temperature plus a fixed seed produces near-identical outputs on repeated runs. For tasks where consistent scores matter, this turns a sample into a measurement.
Seeing cost and quality together changes the decision
Running three separate experiments (one for quality, one for cost, one for latency) produces three snapshots that cannot be compared. Seeing all three simultaneously, across all models, on the same prompt, is what changes the model selection question from a guess to a decision.
These runs were produced with LLM Prover: one prompt, all models, scored simultaneously.
4-run average quality score by model
4-run averages on a root cause classification task. Kimi K3 timed out on all four runs and is excluded.
The $0.60/mo model and the $323/mo model are both correct answers, for different tasks. The benchmark tells you which is which.
Measure your pipeline, not just your models
One prompt, all models, scored simultaneously. Quality, cost, and latency in a single run.
See pricingWhere to start
The right starting point depends on what the stack is showing.
| If you are seeing this | Start here |
|---|---|
| Costs growing, unclear why | Prompt and system prompt audit: run the four-cell benchmark |
| Outputs inconsistent | Run a pipeline diagnostic |
| Model choice feels like a guess | Run the comparison on your actual task |
| Stack works but feels expensive | Check whether a cheaper model now qualifies |
The stack is never done because the landscape keeps changing. Models update, pricing shifts, new options enter. Each pass through the improvement loop lowers the cost floor.
For the full framework covering all 14 levers, see ROI Maxxed AI.