Skip to content

Measure the quality and cost drivers in your AI pipelines

· 6 min read

An AI pipeline is a series of steps where an agent (or a script) calls one or more language models to get work done:

  • summarise a document
  • triage a ticket
  • draft a reply
  • review code

At each step, something chose:

  • which model to call
  • what prompt to send
  • what settings to use

Those choices determine how much the pipeline costs and how good its output is.

When those choices go unmeasured, the pipeline keeps running, but nobody can say whether it is running well or just running.

Three points are where measurement has the most leverage: model selection at build time, output scoring at runtime, and drift detection in production. Each maps to a tool call an agent can make inside the workflow it already runs.

Build time: pick the model from data, not a default

Before committing to a model, an agent can run the actual task against candidates and compare the result.

In a real run on a meeting-note extraction task (parse action items into JSON), four models all returned the correct output. The difference was cost.

ModelLatencyCostClean JSON
GPT-4o Mini2.8s$0.00007yes
DeepSeek V4 Flash5.0s$0.00008yes
Claude Haiku 4.51.4s$0.00065wrapped in code fences
GPT-4.12.9s$0.00086yes

Same correct output, 12x cost spread

The cheapest model and the most expensive both extracted the same three action items. On this task, paying more bought nothing. The only way to know that is to run the comparison on the actual prompt, and an agent can do it before it commits to a model.

Agent
> run_comparison (task prompt, 4 models)
all 4 returned correct JSON | cost: $0.00007 to $0.00086
same output, 12x cost spread. cheapest model wins this task.

The pattern scales: swap the task, keep the process. A coding agent picks a model for code review. A support agent picks one for ticket triage. The comparison is the same tool call; the task is what changes.

A deeper walkthrough on finding the cheapest model that clears a quality bar is next in this series.

Runtime: score the output before it ships

Say your pipeline drafts a customer reply from a support ticket. The reply looks fine to a human, but “fine” is not a number a pipeline can branch on. A rubric turns it into one.

A rubric is a short list of criteria. For a support reply, that might be:

  • Does it address the customer’s actual question?
  • Is the tone professional?
  • Is the answer factually correct given the context?

The agent sends the draft and the rubric to run_evaluation. The judge scores each criterion and returns a breakdown.

CriterionScore
Addresses the question1.0
Professional tone1.0
Factually correct0.0

Overall: 67. The reply read well but cited a policy that does not exist. The judge flagged why: the model hallucinated a refund window the context never mentioned.

That 67 is what the pipeline acts on. Above 90: ship. Below: route back for a redraft with the judge’s reasoning attached, so the next attempt knows what went wrong. The scoring call is async (fire, poll, read), fitting the same pattern as any other tool call in the loop.

The optimization loop post shows this scoring step used iteratively: an agent improving a system prompt one change at a time, each informed by exactly where the previous version lost points.

Start measuring what your pipeline spends and ships

MCP access and the full recipe catalog are on Pro and above.

Get started on Pro

Production: catch drift before users do

The model that scored well last month may not score well today. Model providers update weights, adjust pricing, and change infrastructure without notice. A model your pipeline depends on can:

  • start producing worse output
  • cost more per call
  • respond slower

Nothing in the pipeline will tell you unless you are measuring.

A benchmark is a saved test: the same prompt, the same models, the same scoring rubric, run on a schedule. Each run produces a score per model. When a score drops or a cost spikes compared to the previous run, the system raises a flag.

Here is what that looks like from the agent’s side. The agent picks up where it left off (it saved its place in a durable note last session), checks the latest run, and reports the change.

Agent
> get_agent_notes (recipe_id)
standing note found | suite_id, baseline_run_id
> list_benchmark_runs (suite_id)
run 6 complete | anomaly_flags: score_drop (gpt-4o-mini)
> diff_runs (baseline, run_6)
gpt-4o-mini: quality -8pts, cost +12% | others: stable
regression confirmed on one model. flag for review.

In the chart below, GPT-4o Mini held steady for three runs, then dropped 8 points at run 4 and stayed lower. The other two models were stable throughout. That is the kind of silent regression a scheduled benchmark catches.

Quality trend across scheduled runs, one model regresses at run 4

The cross-session continuity is the part that makes this practical. The agent does not need to be told what it was watching. It finds the note, reads the state, and picks up the monitoring series from where it left off. A closer look at how that memory works is coming in this series.

Raw tools, no lock-in

None of this needs a new framework. LLM Prover exposes its operations as MCP tools, the open standard that agent clients like Cursor, Claude Desktop, and Amazon Q already speak. Point your client at the server and the tools show up; your agent calls them the same way it calls any other tool. No SDK to install, no library to pin a version of. For pipelines that are plain scripts rather than agents, the same operations are available over a REST API. Either way, the only thing you depend on is the tool calls themselves.

Where to start

The fastest path is a five-minute comparison (the tutorial is next in this series). The full catalog of ready-to-paste recipes is in the MCP Recipes Guide. The agentic optimization landing page has the complete picture.

What’s next

mcp evaluation benchmarking agentic