Skip to content

Your LLM stack costs what it costs because you haven't measured it yet.

· 5 min read

The spend is recoverable. Most of it is not where teams look first. Three of the seven models in the benchmarks below did not exist twelve months ago, and model pricing has shifted significantly in that time. A stack that was well-optimised then may not be now. The path forward is measurement, and the first thing measurement reveals is that the model is rarely the whole story.

The compounding effect

Pipeline improvements do not just reduce cost on the current model. They change which models are eligible.

A model that cannot clear your quality threshold with a vague prompt and a loose system prompt may clear it once both are tightened, at a fraction of the cost. That is not a one-time saving. Each pass lowers the cost floor and opens the next optimisation.

The improvement loop has five steps:

  1. Measure: benchmark quality, cost, and latency on your actual task
  2. Tighten inputs: prompt, system prompt, inference parameters
  3. Qualify cheaper models: re-run the benchmark with the tightened inputs
  4. Right-size the architecture: match each task to its minimal model
  5. Monitor for drift: catch regressions before they reach production

The rest of this post is the evidence for each step.

Input precision is a cost lever

Prompt and system prompt are the same lever. Precise inputs produce shorter, more constrained outputs: fewer tokens, lower cost, and better scores. The switcher below shows all four combinations.

In the bottom-right cell, GPT-4.1 Nano and Claude Opus 5 both scored 100. Nano at $0.13/mo, Claude at $10.23/mo. Same score, 79x cost difference.

Claude Opus 5 -- quality score by input precision

Claude Opus 5 on the same task: vague inputs at $0.01884/call, precise inputs at $0.00341/call. A 5.5x cost reduction. Scores went from 0 to 100.

The 5.5x cost reduction is not a system prompt finding specifically. Precise inputs produced shorter outputs. That is the mechanism. Vagueness is not free.

Insight: The cheapest model that clears your quality bar is the right model. Finding it requires measuring the pipeline first.

Parameters are deliberate levers, not defaults

Most teams leave inference parameters at provider defaults. That means the benchmark does not reflect production, and the cost is not being controlled.

Three parameters worth setting deliberately:

  • Max tokens. Caps output length directly. Set it to match your production cap and the benchmark predicts production cost accurately. A task with a predictable output length does not need 1000 tokens of headroom.
  • Prompt caching. If your benchmark uses a fixed system prompt across many runs, caching reduces repeated prefix costs on supported providers. Set it once and it applies to every run.
  • Temperature and seed. Low temperature plus a fixed seed produces near-identical outputs on repeated runs. For tasks where consistent scores matter, this turns a sample into a measurement.
Insight: A benchmark run at production settings is a measurement. A benchmark run at defaults is a guess.

Seeing cost and quality together changes the decision

Running three separate experiments (one for quality, one for cost, one for latency) produces three snapshots that cannot be compared. Seeing all three simultaneously, across all models, on the same prompt, is what changes the model selection question from a guess to a decision.

These runs were produced with LLM Prover: one prompt, all models, scored simultaneously.

4-run average quality score by model

4-run averages on a root cause classification task. Kimi K3 timed out on all four runs and is excluded.

The $0.60/mo model and the $323/mo model are both correct answers, for different tasks. The benchmark tells you which is which.

Insight: The question is not which model is best. It is which model is best for this task.

Measure your pipeline, not just your models

One prompt, all models, scored simultaneously. Quality, cost, and latency in a single run.

See pricing

Where to start

The right starting point depends on what the stack is showing.

If you are seeing thisStart here
Costs growing, unclear whyPrompt and system prompt audit: run the four-cell benchmark
Outputs inconsistentRun a pipeline diagnostic
Model choice feels like a guessRun the comparison on your actual task
Stack works but feels expensiveCheck whether a cheaper model now qualifies

The stack is never done because the landscape keeps changing. Models update, pricing shifts, new options enter. Each pass through the improvement loop lowers the cost floor.

For the full framework covering all 14 levers, see ROI Maxxed AI.

cost-optimisation benchmarking prompt-optimisation roi-maxxed-ai