Skip to content

MCP Reference

MCP Quickstart

Connect an AI agent to LLM Prover via MCP and make your first tool call in under 5 minutes.

What the MCP server gives you

LLM Prover exposes all its core operations as MCP tools. Any MCP-compatible agent or IDE – Amazon Q, Claude Desktop, Cursor, VS Code Copilot – can call them directly. The agent fires a prompt at multiple models, polls for the result, and reads the structured output, all without leaving the chat.

The MCP server is live at the endpoint shown on the MCP page in your dashboard.


Step 1 – Get a developer key

Go to Developer > Keys in your dashboard and create a key. Copy it – it is shown once only.


Step 2 – Connect your client

See the Clients tab on the MCP page for copy-paste config for Amazon Q, Claude Desktop, Cursor, and VS Code. Each client needs the MCP endpoint URL and your developer key.


Step 3 – Discover available tools

Once connected, ask your agent to list the available tools:

“List the LLM Prover MCP tools.”

The server returns 22 tools covering comparisons, evaluations, benchmarks, RAG stores, and files. The full list is on the Tools tab of the MCP page.


Step 4 – Run a comparison

Ask your agent:

“Run a comparison with the prompt ‘Explain the difference between precision and recall in one sentence.’ Use GPT-4.1 Nano and Grok 4.5. Temperature 0.9.”

The agent calls run_comparison, receives a job_id, polls poll_job every 2 seconds until complete, then calls get_comparison to retrieve the result.


The async pattern

Three tools are async – run_comparison, run_evaluation, and trigger_benchmark_run. All three follow the same pattern:

  1. Call the tool. Receive a job_id immediately.
  2. Call poll_job with that job_id. Repeat every 2 seconds.
  3. When status is complete, the full result is in the poll response. No extra call needed.
  4. Stop after 5 minutes if still not complete.

The agent handles this automatically when instructed to poll. Tell it: “Poll every 2 seconds until complete.”


What’s next

  • MCP Tools Reference – every tool, every field, required vs optional
  • MCP Examples – end-to-end worked examples for comparisons, evaluations, and benchmarks