Skip to content

MCP Recipes

pro+ only

Compare Models Over MCP

Fire one prompt at several models through the LLM Prover MCP server, wait for the run to finish, and read the structured comparison.

Run it -- paste into your MCP agent

Highlighted parts are placeholders -- replace them with your own values before (or after) copying.

Using the LLM Prover MCP tools, help me run a model comparison:
1. List the models available to my account and show them to me.
2. Ask me which of those models I want to compare.
3. Before running, tell me my plan's maximum models-per-comparison limit and do
   not exceed it -- if I ask for too many, say so and let me trim the list.
4. Run the comparison on this prompt:
   your prompt here
5. Poll until the run finishes, then give me a short side-by-side summary of each
   model's answer, latency, and cost.

Goal

Run a single prompt against several models through the MCP server and collect the structured per-model result, without leaving the agent chat.

When to use

Reach for this recipe when you want an apples-to-apples look at how different models answer the same prompt – choosing a model for a task, spot-checking a regression, or building a short-list before a full benchmark. It is the fastest path from “I have a prompt” to “here is how five models handled it.”

Steps

  1. Discover the tool. Ask your agent to list the MCP tools and confirm run_comparison and poll_job are present.
  2. Start the run. Call run_comparison with the prompt text and the model IDs you want to compare. The tool returns a run (job) ID immediately – it does not block.
  3. Poll for completion. Call poll_job with the returned ID. Repeat on a short interval until the status comes back as done (or failed).
  4. Read the result. On completion, the job payload carries one row per model: response text, latency, and cost. Summarise or diff them as the task requires.

Verification

  • run_comparison returned a non-empty run ID.
  • poll_job eventually reported done.
  • The result contains one row per requested model, each with response text and a cost figure.

Recipes that lead here