Skip to content

Reference

MCP Recipes Guide

Hand your AI agent a ready-to-paste playbook that runs LLM Prover tools in order, polls for results, and reports back in plain language.

Your AI agent can run a model comparison, score an answer, and watch a benchmark for drift. It does this from inside the client you already work in, calling LLM Prover’s tools directly. A recipe is how you point it at the right job.

What a recipe is

A recipe is a ready-to-paste instruction set you hand your agent. You copy the agent prompt from a recipe page, paste it into Cursor, Claude Desktop, Amazon Q, or Kiro, and the agent runs the LLM Prover tools in the right order, waits for async runs to finish, and reports back in plain language.

The recipe carries the choreography so you do not have to. Here is the full agent prompt from “Compare Models Over MCP”, exactly as it ships.

Using the LLM Prover MCP tools, help me run a model comparison:

  1. List the models available to my account and show them to me.
  2. Ask me which of those models I want to compare.
  3. Before running, tell me my plan’s maximum models-per-comparison limit and do not exceed it. If I ask for too many, say so and let me trim the list.
  4. Run the comparison on this prompt: {your prompt here}
  5. Poll until the run finishes, then give me a short side-by-side summary of each model’s answer, latency, and cost.

You fill in one prompt. The agent does the rest.

Your first run: connect, then compare

The on-ramp is two recipes. “Connect an MCP Client” points your agent at the server and confirms the tools load. “Compare Models Over MCP” is the first real result. Here is a comparison that ran through this exact recipe: one prompt, nine models across five providers, scored side by side.

Agent
> list_models
9 models selected across 5 providers
> run_comparison { prompt, models }
job_id: 13316497 (status: pending)
> poll_job (job_id)
polling... status: complete
9 of 9 models returned. Reading per-model latency, cost, and response.

The agent lists your models, fires the prompt at all nine in parallel, polls the job until it completes, then reads back the per-model result. The async polling is the part you never see. The agent holds the loop and returns when there is something to report.

The result is one row per model.

ModelLatencyCostFollowed the format
GPT-5.14.3s$0.00164yes
Claude Sonnet 53.1s$0.00235added a preamble
GPT-4o Mini3.6s$0.00007yes
Grok 4.79.8s$0.00327yes
DeepSeek V4 Flash12.4s$0.00018yes

Price is no guide to who followed the instruction

The prompt asked for exactly three sentences ending in a one-line analogy.

Insight: Two of the priciest models in the run missed the three-sentence instruction while cheaper models followed it cleanly. Spend buys capability, not obedience. Scoring the output is the only way to know which model did what you asked.

Keep the thread across sessions

Some jobs outlast a single chat. A drift monitor watches a benchmark for weeks. A long run finishes after you have closed your laptop. Several recipes use an agent-state layer so the work survives the session.

The agent keeps a durable note tagged with a recipe ID. The note records which run it cares about and why. When you open a fresh session and ask “any drift?”, the agent finds the note, reads the latest run, and reconciles against the baseline it saved. You never hand it a job ID. It picks up where it left off.

One recipe, “Reconcile Your Runs on Startup”, turns this into a standing habit. Your agent checks what it is watching at the start of every session and reports anything that finished while you were away.

Remember: A standing instruction changes how your agent behaves at every session start, so the recipe shows you the exact text and where it saves it, then waits for your yes. It writes nothing until you approve.

Run these recipes on your own prompts

MCP access and the full recipe catalog are included on Pro and above.

See Pro

The recipe catalog

The recipes group into five tracks, from first connection to production monitoring.

First steps. Connect your client, run your first comparison, score an answer against a known-good response, add your own context to a run, and review recent usage and spend. This is the zero-to-value track.

Make the score mean something. Turn a plain-language description of “good” into a reusable rubric. Run a single-variable A/B test on your prompt or your system prompt. Ask for the cheapest model that clears a quality bar and get a recommendation, not a table.

Set and forget. Stand up a drift monitor on a schedule, set plain-language alerts, track a long run across sessions, and have each new session reconcile automatically. This is the production track, and most of it leans on the agent-state layer.

Context and compliance. Create a RAG store and smoke-test retrieval before you rely on it. Check a response against your regulations clause by clause, with the regulatory text the judge used. Both are Pro+.

Housekeeping. Tidy stale stores, rubrics, prompts, and runs, with one confirmation per deletion.

Each recipe page states when to use it, what it needs, what it produces, and the exact steps, so you can read it before you run it. Browse them in the Recipes section.

When to reach past recipes

Recipes are the curated front door. They cover the common jobs with a prompt you can paste and tune. Behind them sits the full tool surface, over forty tools spanning comparison, evaluation, benchmarks, rubrics, RAG stores, custom endpoints, and the agent-state layer.

When you own your own agent orchestration, that surface is yours to wire directly. LLM Prover becomes the measurement node inside your own loops and graphs, the place where a model decision gets scored before it ships.

Before you start

Two things are worth knowing up front.

MCP access starts at Pro. The RAG store and compliance recipes need Pro+. The Starter plan does not include MCP.

Uploading files is a dashboard step. MCP creates the RAG store and queries it, but the file upload itself happens in the dashboard. The RAG setup recipe smoke-tests retrieval afterward, so a store that failed to ingest shows up before you run a real query against it. Recipes are written to tell you when something is not ready rather than return a confident answer built on broken context.

What’s next