Skip to content

From connected to running evals in under a minute

· 3 min read

If you already run an MCP client (Cursor, Claude Desktop, Amazon Q, Kiro), the distance between where you are now and running real model evaluations is short. This post times it.

Connecting is a paste, not a project

The “Connect an MCP Client” recipe handles setup. You copy its instruction block into your client, point it at the LLM Prover MCP server, and the tools appear in your agent’s toolset. There is no SDK to install and no integration code to write. The full walkthrough is in the MCP Recipes Guide. Once the tools load, the clock below starts.

The clock starts

The task: summarise the three biggest risks in a contract clause, under 50 words, across three models. The agent fired the comparison, polled until it finished, and had all three responses back.

Agent
> run_comparison (clause-risk prompt, 3 models)
job_id: bab1e605 status: pending
> poll_job
polling... complete (2.7s)
claude-haiku-4-5 | gpt-4o-mini | gpt-4.1 all returned
3 responses in. one of them misread the clause. the pipeline flagged it.

Start to finish, including the time to read the result back, this took about 43 seconds. The comparison itself ran in 2.7 seconds. Three models, one prompt, no dashboard.

ModelLatencyCost
Claude Haiku 4.51.7s$0.00056
GPT-4o Mini2.6s$0.00004
GPT-4.12.7s$0.00046

Three summaries of the same clause, back in under three seconds

The fast answer was wrong

GPT-4o Mini is the cheapest model in the run and returned a clean, confident summary. It was also wrong. The clause caps the Customer’s liability at twelve months of fees. GPT-4o Mini reported “unlimited liability for the Supplier,” inverting who the cap protects.

The diagnostic layer caught it without anyone reading the clause line by line. It flagged the misread as a high-impact finding: the model stated unlimited Supplier liability when the text caps Customer liability. The other two models read the clause correctly.

Warning: The cheapest, fastest response was also the one that got the clause backwards. Speed and price tell you nothing about whether an answer is correct. A comparison shows you what each model said; scoring is what tells you which one to trust.
Insight: A model misread a contract clause and the pipeline caught it in under a minute. That is a real production problem, found and flagged before it reached a human, in less time than it takes to read the clause yourself.

Run your first comparison

Connect an MCP client and compare models on your own prompt. Pro and above.

Get started on Pro

Where this goes next

A comparison gets you three answers in seconds. Turning that into a decision (“which model should this pipeline actually use?”) means scoring the responses against a rubric that defines what good looks like for your task. That is the full walkthrough in How to find the right LLM for your use case, where a rubric and a benchmark separate a field of models that all looked fine at first glance.

For the bigger picture of where measurement fits across an AI pipeline, start with Measure the quality and cost drivers in your AI pipelines.

What’s next

mcp comparison agentic tutorial