Reference
MCP Recipes Guide
Hand your AI agent a ready-to-paste playbook that runs LLM Prover tools in order, polls for results, and reports back in plain language.
Your AI agent can run a model comparison, score an answer, and watch a benchmark for drift. It does this from inside the client you already work in, calling LLM Prover’s tools directly. A recipe is how you point it at the right job.
What a recipe is
A recipe is a ready-to-paste instruction set you hand your agent. You copy the agent prompt from a recipe page, paste it into Cursor, Claude Desktop, Amazon Q, or Kiro, and the agent runs the LLM Prover tools in the right order, waits for async runs to finish, and reports back in plain language.
The recipe carries the choreography so you do not have to. Here is the full agent prompt from “Compare Models Over MCP”, exactly as it ships.
Using the LLM Prover MCP tools, help me run a model comparison:
- List the models available to my account and show them to me.
- Ask me which of those models I want to compare.
- Before running, tell me my plan’s maximum models-per-comparison limit and do not exceed it. If I ask for too many, say so and let me trim the list.
- Run the comparison on this prompt: {your prompt here}
- Poll until the run finishes, then give me a short side-by-side summary of each model’s answer, latency, and cost.
You fill in one prompt. The agent does the rest.
Your first run: connect, then compare
The on-ramp is two recipes. “Connect an MCP Client” points your agent at the server and confirms the tools load. “Compare Models Over MCP” is the first real result. Here is a comparison that ran through this exact recipe: one prompt, nine models across five providers, scored side by side.
The agent lists your models, fires the prompt at all nine in parallel, polls the job until it completes, then reads back the per-model result. The async polling is the part you never see. The agent holds the loop and returns when there is something to report.
The result is one row per model.
| Model | Latency | Cost | Followed the format |
|---|---|---|---|
| GPT-5.1 | 4.3s | $0.00164 | yes |
| Claude Sonnet 5 | 3.1s | $0.00235 | added a preamble |
| GPT-4o Mini | 3.6s | $0.00007 | yes |
| Grok 4.7 | 9.8s | $0.00327 | yes |
| DeepSeek V4 Flash | 12.4s | $0.00018 | yes |
Price is no guide to who followed the instruction
The prompt asked for exactly three sentences ending in a one-line analogy.
Keep the thread across sessions
Some jobs outlast a single chat. A drift monitor watches a benchmark for weeks. A long run finishes after you have closed your laptop. Several recipes use an agent-state layer so the work survives the session.
The agent keeps a durable note tagged with a recipe ID. The note records which run it cares about and why. When you open a fresh session and ask “any drift?”, the agent finds the note, reads the latest run, and reconciles against the baseline it saved. You never hand it a job ID. It picks up where it left off.
One recipe, “Reconcile Your Runs on Startup”, turns this into a standing habit. Your agent checks what it is watching at the start of every session and reports anything that finished while you were away.
Run these recipes on your own prompts
MCP access and the full recipe catalog are included on Pro and above.
See ProThe recipe catalog
The recipes group into five tracks, from first connection to production monitoring.
First steps. Connect your client, run your first comparison, score an answer against a known-good response, add your own context to a run, and review recent usage and spend. This is the zero-to-value track.
Make the score mean something. Turn a plain-language description of “good” into a reusable rubric. Run a single-variable A/B test on your prompt or your system prompt. Ask for the cheapest model that clears a quality bar and get a recommendation, not a table.
Set and forget. Stand up a drift monitor on a schedule, set plain-language alerts, track a long run across sessions, and have each new session reconcile automatically. This is the production track, and most of it leans on the agent-state layer.
Context and compliance. Create a RAG store and smoke-test retrieval before you rely on it. Check a response against your regulations clause by clause, with the regulatory text the judge used. Both are Pro+.
Housekeeping. Tidy stale stores, rubrics, prompts, and runs, with one confirmation per deletion.
Each recipe page states when to use it, what it needs, what it produces, and the exact steps, so you can read it before you run it. Browse them in the Recipes section.
When to reach past recipes
Recipes are the curated front door. They cover the common jobs with a prompt you can paste and tune. Behind them sits the full tool surface, over forty tools spanning comparison, evaluation, benchmarks, rubrics, RAG stores, custom endpoints, and the agent-state layer.
When you own your own agent orchestration, that surface is yours to wire directly. LLM Prover becomes the measurement node inside your own loops and graphs, the place where a model decision gets scored before it ships.
Before you start
Two things are worth knowing up front.
MCP access starts at Pro. The RAG store and compliance recipes need Pro+. The Starter plan does not include MCP.
Uploading files is a dashboard step. MCP creates the RAG store and queries it, but the file upload itself happens in the dashboard. The RAG setup recipe smoke-tests retrieval afterward, so a store that failed to ingest shows up before you run a real query against it. Recipes are written to tell you when something is not ready rather than return a confident answer built on broken context.
What’s next
MCP Quickstart
Connect an AI agent to LLM Prover via MCP and make your first tool call in under 5 minutes.
MCP Tools Reference
Every LLM Prover MCP tool, its input fields, and what it returns.
Monitoring Guide
How LLM Prover detects changes in model quality, cost, and latency, and how to configure alerts so you know before users do.