Your AI agents can now benchmark and monitor other AI models
A marketing agent running inside an IDE called run_comparison, polled for results, retrieved response text, and made a qualitative judgment about which model produced the sharpest answer. No human touched the app. That happened today, using LLM Prover’s new MCP server. This is what agent-native tooling looks like in practice.
What shipped
MCP (Model Context Protocol) is the standard that lets AI agents discover and call tools in their environment. LLM Prover now exposes a full MCP server with 22 tools covering: