Measure the quality and cost drivers in your AI pipelines
An AI pipeline is a series of steps where an agent (or a script) calls one or more language models to get work done:
- summarise a document
- triage a ticket
- draft a reply
- review code
At each step, something chose:
- which model to call
- what prompt to send
- what settings to use
Those choices determine how much the pipeline costs and how good its output is.
When those choices go unmeasured, the pipeline keeps running, but nobody can say whether it is running well or just running.
Three points are where measurement has the most leverage: model selection at build time, output scoring at runtime, and drift detection in production. Each maps to a tool call an agent can make inside the workflow it already runs.
Build time: pick the model from data, not a default
Before committing to a model, an agent can run the actual task against candidates and compare the result.
In a real run on a meeting-note extraction task (parse action items into JSON), four models all returned the correct output. The difference was cost.
| Model | Latency | Cost | Clean JSON |
|---|---|---|---|
| GPT-4o Mini | 2.8s | $0.00007 | yes |
| DeepSeek V4 Flash | 5.0s | $0.00008 | yes |
| Claude Haiku 4.5 | 1.4s | $0.00065 | wrapped in code fences |
| GPT-4.1 | 2.9s | $0.00086 | yes |
Same correct output, 12x cost spread
The cheapest model and the most expensive both extracted the same three action items. On this task, paying more bought nothing. The only way to know that is to run the comparison on the actual prompt, and an agent can do it before it commits to a model.
The pattern scales: swap the task, keep the process. A coding agent picks a model for code review. A support agent picks one for ticket triage. The comparison is the same tool call; the task is what changes.
A deeper walkthrough on finding the cheapest model that clears a quality bar is next in this series.
Runtime: score the output before it ships
Say your pipeline drafts a customer reply from a support ticket. The reply looks fine to a human, but “fine” is not a number a pipeline can branch on. A rubric turns it into one.
A rubric is a short list of criteria. For a support reply, that might be:
- Does it address the customer’s actual question?
- Is the tone professional?
- Is the answer factually correct given the context?
The agent sends the draft and the rubric to run_evaluation. The judge scores each criterion and returns a breakdown.
| Criterion | Score |
|---|---|
| Addresses the question | 1.0 |
| Professional tone | 1.0 |
| Factually correct | 0.0 |
Overall: 67. The reply read well but cited a policy that does not exist. The judge flagged why: the model hallucinated a refund window the context never mentioned.
That 67 is what the pipeline acts on. Above 90: ship. Below: route back for a redraft with the judge’s reasoning attached, so the next attempt knows what went wrong. The scoring call is async (fire, poll, read), fitting the same pattern as any other tool call in the loop.
The optimization loop post shows this scoring step used iteratively: an agent improving a system prompt one change at a time, each informed by exactly where the previous version lost points.
Start measuring what your pipeline spends and ships
MCP access and the full recipe catalog are on Pro and above.
Production: catch drift before users do
The model that scored well last month may not score well today. Model providers update weights, adjust pricing, and change infrastructure without notice. A model your pipeline depends on can:
- start producing worse output
- cost more per call
- respond slower
Nothing in the pipeline will tell you unless you are measuring.
A benchmark is a saved test: the same prompt, the same models, the same scoring rubric, run on a schedule. Each run produces a score per model. When a score drops or a cost spikes compared to the previous run, the system raises a flag.
Here is what that looks like from the agent’s side. The agent picks up where it left off (it saved its place in a durable note last session), checks the latest run, and reports the change.
In the chart below, GPT-4o Mini held steady for three runs, then dropped 8 points at run 4 and stayed lower. The other two models were stable throughout. That is the kind of silent regression a scheduled benchmark catches.
Quality trend across scheduled runs, one model regresses at run 4
The cross-session continuity is the part that makes this practical. The agent does not need to be told what it was watching. It finds the note, reads the state, and picks up the monitoring series from where it left off. A closer look at how that memory works is coming in this series.
Raw tools, no lock-in
None of this needs a new framework. LLM Prover exposes its operations as MCP tools, the open standard that agent clients like Cursor, Claude Desktop, and Amazon Q already speak. Point your client at the server and the tools show up; your agent calls them the same way it calls any other tool. No SDK to install, no library to pin a version of. For pipelines that are plain scripts rather than agents, the same operations are available over a REST API. Either way, the only thing you depend on is the tool calls themselves.
Where to start
The fastest path is a five-minute comparison (the tutorial is next in this series). The full catalog of ready-to-paste recipes is in the MCP Recipes Guide. The agentic optimization landing page has the complete picture.
What’s next
MCP Recipes Guide
Hand your AI agent a ready-to-paste playbook that runs LLM Prover tools in order, polls for results, and reports back in plain language.
Monitoring Guide
How LLM Prover detects changes in model quality, cost, and latency, and how to configure alerts so you know before users do.
Scoring Guide
How LLM Prover calculates scores, what affects them, and how to use the diagnostic judge to find out why.