Your AI agents can now benchmark and monitor other AI models
A marketing agent running inside an IDE called run_comparison, polled for results, retrieved response text, and made a qualitative judgment about which model produced the sharpest answer. No human touched the app. That happened today, using LLM Prover’s new MCP server. This is what agent-native tooling looks like in practice.
What shipped
MCP (Model Context Protocol) is the standard that lets AI agents discover and call tools in their environment. LLM Prover now exposes a full MCP server with 22 tools covering:
- Comparisons
- Evaluations
- Benchmarks
- RAG stores
- Model discovery
Any MCP-capable client discovers them automatically. Claude Desktop, Cursor, VS Code Copilot, and Amazon Q all work without integration code or API wiring.
The tools are async-native: fire a run, get a job_id back immediately, poll until complete, retrieve the full result. That pattern is what makes this usable inside a real agent loop rather than a blocking call that stalls the workflow.
If you are building a pipeline rather than using an agent IDE, the same operations are available via REST with a developer key. Same async pattern, same results, same 22 endpoints.
Full tool reference and client setup guides are in the MCP and API docs.
The run that proved it
The agent fired a comparison with these settings:
Prompt: In one sentence, explain why agent-first SaaS will be bigger than human-oriented SaaS.
Models: Grok 4.5, Claude Sonnet 4.5, GPT-4.1 Nano, DeepSeek V4 Flash, Qwen 3.8 27B
Temperature: 0.9. No system prompt. No context store.
The agent polled until complete, retrieved the full result, and assessed the responses, all within a single IDE conversation. Here is what came back:
The cost spread across those five responses:
Cost per call -- agent-first SaaS comparison run
DeepSeek produced the sharpest framing. “Zero marginal labor costs” and “capped by the time, speed, and cost of the human operators it serves” is a more precise argument than the others. GPT-4.1 Nano was the weakest on this task, but at $0.000018 per call it is 75x cheaper than Grok 4.5. Qwen 3.8 27B returned in 920ms, fastest by a wide margin.
The agent made that assessment. Not a human reviewing a dashboard.
What agents can do with this
Agentic eval loops. An agent writes a prompt, evaluates it with run_evaluation, reads the diagnostic findings, refines the prompt, and re-evaluates, all without leaving the workflow. The improvement loop that previously required a human to interpret scores and decide what to change runs end to end. There is no deterministic equivalent to this: the agent is making judgment calls at each step.
Model substitution on degradation. When a benchmark run returns a quality drop below threshold, an agent can call run_comparison on candidate replacements, assess the results, and propose a substitution before a human has noticed anything is wrong. That requires judgment, not just threshold logic.
Model selection at build time. An agent building a pipeline can call run_comparison on its candidate models before committing to one. The decision is based on data from the actual task, not a benchmark someone else ran on different prompts.
Drift detection across a fleet. An agent managing multiple deployments can call list_benchmark_runs across all suites, identify which ones have drifted, and surface only the ones that need attention. The monitoring work runs on a schedule; a human is only involved when something actually breaks.
MCP access is on Pro and above. The REST API is available on all tiers. Connect your client in under five minutes from the MCP page in your dashboard.
Add LLM Prover to your agent
MCP on Pro and above. REST API on all tiers. Connect in under five minutes.
Get started on Pro