Your AI agents can now benchmark and monitor other AI models.
Any MCP-capable agent can measure your prompts running against your data, analyse results and make decisions without leaving its workflow.
Get started on ProAgents run comparisons, read the results, and make decisions.
ProA marketing agent running inside an IDE called
run_comparison across five models, polled for results, retrieved the full response text, and assessed which model produced the sharpest answer. No human touched the app. That is what MCP access looks like in practice. Read the full account in the MCP launch post.Cost per call -- agent-first SaaS comparison run
Agents select models at build time, based on the actual task.
ProAn agent building a pipeline calls
run_comparison on its candidate models before committing to one. The decision is based on data from the actual task, not a benchmark someone else ran on different prompts. The agent reads the cost, latency, and response quality, and picks the model that fits. No human reviews a dashboard.Quality vs latency -- bubble size = relative cost
Agents refine their own prompts through an automated eval loop.
ProAn agent writes a prompt, calls
run_evaluation to score it against a rubric, reads the diagnostic findings, refines the prompt, and re-evaluates. Quality climbs across iterations. Cost per call drops as the prompt tightens. The improvement loop that previously required a human to interpret scores and decide what to change runs end to end.Agentic eval loop -- quality and cost per iteration
Agents substitute models before a human notices the drop.
ProWhen a benchmark run returns a quality drop below threshold, an agent calls
run_comparison on candidate replacements, reads the results, and proposes a substitution. The decision is based on data from the actual task. A human reviews the proposal. The agent does the measurement work.Quality score -- model drop detected, substitution applied
Agents detect drift across a fleet, without a dashboard open.
ProAn agent managing multiple deployments calls
list_benchmark_runs across all suites, identifies which ones have drifted, and surfaces only the ones that need attention. The monitoring runs on a schedule. The agent filters the noise. A human sees only what requires a decision.Threshold breaches across benchmark fleet -- weekly
Where to start
MCP access is on Pro and above. Any MCP-capable client, including Claude Desktop, Cursor, VS Code Copilot and Amazon Q, discovers the 22 tools automatically. No integration code required. See the MCP docs.
Starter $49/mo | Pro $129/mo | Pro+ $199/mo | Enterprise $499/mo | |
|---|---|---|---|---|
| MCP server access | -- | ✓ | ✓ | ✓ |
| Models per run | 5 | 8 | 10 | 20 |
| RAG context in runs | -- | 1 store, 20 files | 5 stores, 100 files | unlimited |
| Compliance scoring | -- | -- | ✓ | ✓ |
| Benchmark schedule | -- | daily | 4x daily | hourly |
| History retention | 90 days | 365 days | 365 days | unlimited |
| Developer keys | 1 | 5 | 10 | 50 |
| Credits included | $12 | $45 | $73 | $150 |
Add LLM Prover to your agent.
MCP on Pro and above. Connect your client in under five minutes.