Skip to content

MCP Recipes

Pro

Run With a Preset (or Set One Up)

Teach your agent to resolve an underspecified ask (give this a standard comparison) against your saved presets, and to offer to save a good ad-hoc config as a reusable preset so next time you can just name it.

⚠

Agentic and programmatic use. Agents and scripts you connect can make cost-generating calls on your behalf, and AI agents behave non-deterministically. You are responsible for monitoring and bounding your own automation. See Terms for details.

Run it: paste into your MCP agent

Highlighted parts are placeholders. Replace them with your own values before (or after) copying.

Using the LLM Prover MCP tools, run with a PRESET (a named, saved run configuration), or
help me set one up. A preset is what lets me be terse: I name a task, the preset supplies
the config (models, scoring, params), and you inject only my content. You never guess the
config; a preset supplies it, or I specify it.

1. Confirm the tools you need are available (list the tools): list_presets, run_preset,
   create_preset, create_benchmark_from_preset, plus run_comparison / run_evaluation and
   poll_job.

WHEN I ASK FOR A RUN:
2. If my ask is UNDERSPECIFIED (e.g. "give this a standard comparison", "run my usual eval
   on this", "score this the normal way"), call list_presets FIRST. Match my words to a
   preset by its name and description. If one clearly matches, confirm it in one line (e.g.
   "using your 'standard' preset: 4 models, default rubric") and run it:
     run_preset(preset_id=<the match>, content=<my prompt or source text>).
   Async, so poll_job every 5 seconds to completion, then report the result. I supplied only
   content; everything else came from the preset.
3. If the match is AMBIGUOUS (two plausible presets, or you are unsure), ask me which rather
   than guessing. A wrong preset silently runs the wrong config.
4. If NO preset matches an underspecified ask, do NOT invent a config and do NOT run. Tell
   me plainly there is no matching preset, and ask me to specify the run ONCE: which models,
   what scoring (a gold-standard answer or a rubric, or none), any params. Then run it the
   normal way (run_comparison / run_evaluation).

PROMOTE A GOOD AD-HOC RUN TO A PRESET (conditional; do not over-offer):
5. After a run that I specified by hand, the result carries a `promotable` block. Offer to
   save the config as a preset ONLY when ALL of these hold:
     - promotable.eligible is true (the server's check: a substantive run, not one launched
       from a preset, not a bare default run), AND
     - I specified the config myself, NOT you. If I just said "pick 3 models and run this"
       and YOU chose them, there is no considered config worth saving; do not offer.
   When you do offer: propose a name (ask me, or suggest one like "standard") and show me
   the EXACT config you would save. I may want all of it, or only a SUBSET (e.g. "just the
   models, not the params"). Save only what I approve:
     create_preset(name, run_type, <the approved config subset>).
   NEVER save a preset silently: a preset I did not see becomes a config I unknowingly
   reuse forever. If promotable.eligible is false, or the config was yours not mine, say
   nothing about presets.
6. Do NOT offer to promote a run that was launched FROM a preset (its config is already a
   preset). The `promotable` block will say so; respect it.

SET UP A PRESET DIRECTLY (if I ask to):
7. If I ask to create a preset outright, gather the config with me (models; a run_type of
   performance, quality, or compliance; a context store if the models should read one;
   scoring for quality/compliance; params, system prompt, any subset I care to fix), show
   it back for approval, then create_preset. Write the description for TASK-MATCHING
   ("use for: ...") so a future
   underspecified ask resolves to it cleanly.

MONITOR WITH A PRESET:
8. If I ask to "monitor", "track over time", or "set up a recurring test" for a prompt and a
   preset fits, use create_benchmark_from_preset(preset_id, content, [schedule]) rather than
   building a benchmark config from scratch. Tell me whether it is a quality benchmark (the
   preset has scoring) or a performance one (it does not).

ON ANY FAILURE or clear mismatch (a preset will not run, a referenced model is retired, the
save is rejected): tell me honestly what happened and at which step. A preset that references
a retired model or a model above my tier is blocked with a clear reason. Relay it (heal the
preset, or upgrade) rather than silently falling back to a different config.

Goal

Stop re-specifying the same run. A preset is a named, saved configuration (models, scoring, params, system prompt) that you build once and then invoke by name: “give this prompt a standard comparison” resolves to your saved “standard” preset, and the agent injects only your prompt. The config is the preset’s job, not something the agent guesses.

When to use

Reach for this once you have found a setup you run repeatedly. Save it as a preset and every future run is one sentence. It also covers the reverse: when you specify a good run by hand, the agent offers to save that config so you do not have to type it again.

How it works

  • A preset supplies the config; you supply the content. run_preset takes a preset name and your prompt. Models, scoring, and params come from the stored preset, so an underspecified ask becomes a reproducible, one-sentence run.
  • No preset, no guessing. If you ask for a run the agent cannot resolve to a preset, it does not invent models and scoring and quietly run. It asks you to specify the run once, then offers to save it so next time is terse.
  • Promotion is earned, and never silent. After a run you configured yourself, the agent offers to save that config as a preset, but only when there is something deliberate worth saving (not a bare default run, and not a config the agent picked for you). It shows the exact config, lets you save all of it or a subset, and saves nothing without your approval. A run that already came from a preset is never offered for promotion.
  • Curated to start. A small set of curated presets ships ready to run on any tier, so you can see what a preset does before building your own. Curated presets do not count against your preset limit.
  • Presets stay honest. A preset is re-validated against your current plan and the live model registry every time it runs. If it references a retired model or a model above your tier, the run is blocked with a clear reason rather than silently substituting a different model.

Verification

  • An underspecified ask resolved to a named preset (confirmed before running), or, when none matched, the agent asked for a one-off spec instead of guessing.
  • A preset run injected only your content; models, scoring, and params came from the preset.
  • A promote offer appeared only after a substantive, user-specified ad-hoc run, showed the exact config for approval, and never saved silently; a preset-launched run was not offered for promotion.
  • A preset referencing a retired or above-tier model was blocked with a clear reason, not silently run on a substitute.

Where this leads

Natural next recipes once you are comfortable with this one.