From connected to running evals in under a minute
If you already run an MCP client (Cursor, Claude Desktop, Amazon Q, Kiro), the distance between where you are now and running real model evaluations is short. This post times it.
Connecting is a paste, not a project
The “Connect an MCP Client” recipe handles setup. You copy its instruction block into your client, point it at the LLM Prover MCP server, and the tools appear in your agent’s toolset. There is no SDK to install and no integration code to write. The full walkthrough is in the MCP Recipes Guide. Once the tools load, the clock below starts.
The clock starts
The task: summarise the three biggest risks in a contract clause, under 50 words, across three models. The agent fired the comparison, polled until it finished, and had all three responses back.
Start to finish, including the time to read the result back, this took about 43 seconds. The comparison itself ran in 2.7 seconds. Three models, one prompt, no dashboard.
| Model | Latency | Cost |
|---|---|---|
| Claude Haiku 4.5 | 1.7s | $0.00056 |
| GPT-4o Mini | 2.6s | $0.00004 |
| GPT-4.1 | 2.7s | $0.00046 |
Three summaries of the same clause, back in under three seconds
The fast answer was wrong
GPT-4o Mini is the cheapest model in the run and returned a clean, confident summary. It was also wrong. The clause caps the Customer’s liability at twelve months of fees. GPT-4o Mini reported “unlimited liability for the Supplier,” inverting who the cap protects.
The diagnostic layer caught it without anyone reading the clause line by line. It flagged the misread as a high-impact finding: the model stated unlimited Supplier liability when the text caps Customer liability. The other two models read the clause correctly.
Run your first comparison
Connect an MCP client and compare models on your own prompt. Pro and above.
Where this goes next
A comparison gets you three answers in seconds. Turning that into a decision (“which model should this pipeline actually use?”) means scoring the responses against a rubric that defines what good looks like for your task. That is the full walkthrough in How to find the right LLM for your use case, where a rubric and a benchmark separate a field of models that all looked fine at first glance.
For the bigger picture of where measurement fits across an AI pipeline, start with Measure the quality and cost drivers in your AI pipelines.
What’s next
MCP Recipes Guide
Hand your AI agent a ready-to-paste playbook that runs LLM Prover tools in order, polls for results, and reports back in plain language.
Comparisons Guide
Fire a prompt at multiple models simultaneously and compare cost, latency, and output side by side.
Scoring Guide
How LLM Prover calculates scores, what affects them, and how to use the diagnostic judge to find out why.