Hand your agent a recipe, not an API
LLM Prover’s MCP server already lets any MCP-capable agent call its tools directly. The gap was knowing which tools to call, in what order, and how to handle the async runs in between. Recipes close that gap.
A recipe is a copy-paste instruction set you hand your agent. Copy it from the recipe page, paste it into Cursor, Claude Desktop, Amazon Q, or Kiro, and the agent runs the LLM Prover tools in the right order, waits for runs to finish, and reports back in plain language. You bring the prompt and the agent does the rest.
What shipped
A catalog of recipes, grouped into five tracks that follow the work from first connection to production monitoring.
- First steps. Connect your client, run your first comparison, score an answer against a known-good response, add your own context to a run, review recent spend.
- Make the score mean something. Turn a plain-language description of “good” into a reusable rubric. A/B test a prompt or a system prompt. Ask for the cheapest model above a quality bar.
- Set and forget. Stand up a drift monitor on a schedule, set plain-language alerts, track a long run across sessions.
- Context and compliance. Create a RAG store and smoke-test retrieval. Check a response against your regulations, clause by clause. Both are Pro+.
- Housekeeping. Tidy stale stores, rubrics, prompts, and runs.
Each recipe page states when to use it, what it needs, what it produces, and the exact steps.
How a recipe remembers
The jobs that matter most are the ones that outlast a single chat: a drift monitor that watches a benchmark for weeks, a long run that finishes after you have closed your laptop. For those to work, the agent needs to remember what it was doing after the session ends.
Recipes solve this with durable notes. The note tools are a first-class part of the MCP surface, and the recipe instructs the agent to use them: it writes the loop’s state into a note tagged with a recipe ID, then reads it back to stay honest. The note records which run it is watching, the baseline it compares against, and where it is in a multi-step job.
A drift monitor is the clearest case. Close your laptop, open a fresh session a week later, ask “any drift?”, and the agent finds the note by its recipe ID, reads the latest run, and reconciles against the baseline it saved. You never hand it a job ID. The recipe is the choreography, the note is the memory, and together they turn a one-shot tool call into a job that survives you walking away.
When recipes run out
Recipes are the curated front door. They cover the common jobs with a prompt you can paste and tune. When you own your own agent orchestration, the full tool surface is yours to wire directly, and LLM Prover becomes the measurement node inside your own loops and graphs. That is a longer story for another post.
For now, the fastest path from “I have a prompt” to “my agent is benchmarking models for me” is a recipe you paste. Start with Connect an MCP Client, then Compare Models Over MCP. The full walkthrough is in the MCP Recipes Guide.
Browse the recipe catalog
MCP and the full recipe catalog are on Pro and above. Connect your client in under five minutes.
Get started on Pro