Skip to content

Your AI agents can now benchmark and monitor other AI models.

Any MCP-capable agent can measure your prompts running against your data, analyse results and make decisions without leaving its workflow.

Get started on Pro

Agents run comparisons, read the results, and make decisions.

Pro
A marketing agent running inside an IDE called run_comparison across five models, polled for results, retrieved the full response text, and assessed which model produced the sharpest answer. No human touched the app. That is what MCP access looks like in practice. Read the full account in the MCP launch post.

Cost per call -- agent-first SaaS comparison run

Agents select models at build time, based on the actual task.

Pro
An agent building a pipeline calls run_comparison on its candidate models before committing to one. The decision is based on data from the actual task, not a benchmark someone else ran on different prompts. The agent reads the cost, latency, and response quality, and picks the model that fits. No human reviews a dashboard.

Quality vs latency -- bubble size = relative cost

Agents refine their own prompts through an automated eval loop.

Pro
An agent writes a prompt, calls run_evaluation to score it against a rubric, reads the diagnostic findings, refines the prompt, and re-evaluates. Quality climbs across iterations. Cost per call drops as the prompt tightens. The improvement loop that previously required a human to interpret scores and decide what to change runs end to end.

Agentic eval loop -- quality and cost per iteration

Agents substitute models before a human notices the drop.

Pro
When a benchmark run returns a quality drop below threshold, an agent calls run_comparison on candidate replacements, reads the results, and proposes a substitution. The decision is based on data from the actual task. A human reviews the proposal. The agent does the measurement work.

Quality score -- model drop detected, substitution applied

Agents detect drift across a fleet, without a dashboard open.

Pro
An agent managing multiple deployments calls list_benchmark_runs across all suites, identifies which ones have drifted, and surfaces only the ones that need attention. The monitoring runs on a schedule. The agent filters the noise. A human sees only what requires a decision.

Threshold breaches across benchmark fleet -- weekly

Where to start

MCP access is on Pro and above. Any MCP-capable client, including Claude Desktop, Cursor, VS Code Copilot and Amazon Q, discovers the 22 tools automatically. No integration code required. See the MCP docs.
Starter
$49/mo
Pro
$129/mo
Pro+
$199/mo
Enterprise
$499/mo
MCP server access--✓✓✓
Models per run581020
RAG context in runs--1 store, 20 files5 stores, 100 filesunlimited
Compliance scoring----✓✓
Benchmark schedule--daily4x dailyhourly
History retention90 days365 days365 daysunlimited
Developer keys151050
Credits included$12$45$73$150

Add LLM Prover to your agent.

MCP on Pro and above. Connect your client in under five minutes.

Get started on Pro