Skip to content

MCP Recipes

Pro

A/B Test Your System Prompt

Save two or more system prompts, run the same task through each on one model, score them the same way, and see which set of instructions produces the best result.

Run it: paste into your MCP agent

Highlighted parts are placeholders. Replace them with your own values before (or after) copying.

Using the LLM Prover MCP tools, run a single-variable A/B test on the SYSTEM PROMPT: same
model, same task prompt, same scoring, only the system prompt changes.

1. Confirm the tools you need are available (list the tools).
2. Ask me for the system prompt variants to test (two or more) and the task prompt they
   will all run against, if I have not given them.
3. Lock the controls so this is a fair test: pick ONE model (list_models; sensible default
   or ask me) and ONE scoring method -- a gold-standard answer (expected_output) or a saved
   rubric (rubric_id). Model, task prompt, and scoring stay identical across variants.
4. Save each system prompt variant: create_system_prompt(name, content) -- note each
   prompt_id. (Saving them makes the winner reusable afterward, and keeps the test clean.)
   Respect my plan's system-prompts cap; if I am near it, say so.
5. Run each variant as its own scored evaluation, identical except for the system prompt:
   for each, run_evaluation(prompt=<the task prompt>, models=<the one model>,
   system_prompt_id=<that variant's id>, expected_output=<gold> OR rubric_id=<rubric>).
   Each is async -- poll_job every 5 seconds to completion (stop after 5 minutes).
6. Compare scores across variants and recommend the best system prompt. Show each variant's
   quality_score side by side with the judge's reasoning and name the winner; if two are
   within a few points, call it close rather than overstating a small difference.
7. Hand back the winning system prompt's prompt_id and remind me it is saved and reusable
   (attach it via system_prompt_id in any comparison, evaluation, or benchmark). Offer once
   to monitor the winning setup over time, if it fits.

ON ANY FAILURE or clear mismatch (a create fails, a variant will not score): tell me
honestly what happened and at which step -- compare the variants that succeeded, name the
ones excluded, and do not invent a score. If some system prompts were saved before a
failure, tell me their prompt_ids so we do not create duplicates on retry.

Goal

Find the system prompt that gets the best behaviour out of a model for your task. You give several sets of instructions; the agent saves each, runs the same task under each on the same model, scores them the same way, and tells you which system prompt wins – with the winner left saved and reusable.

When to use

Reach for this when the model’s behaviour depends on how you instruct it and you want to choose the best instructions with evidence. It is the system-prompt-stage A/B test: fix the model, the task, and the scoring; vary only the system prompt.

How it works

  • One variable. Model, task prompt, and scoring are fixed; only the system prompt changes, so score differences are attributable to the instructions.
  • Variants are saved, not throwaway. Each system prompt is saved via create_system_prompt, so the winner is immediately reusable (and the test references stable prompt_ids).
  • Scored and honest. Each variant gets a quality_score and reasoning; close results are called close, not spun into a winner.

Verification

  • Every variant used the same model, task prompt, and scoring – only the system prompt differed.
  • Each system prompt was saved (has a prompt_id) and has its own score.
  • The recommendation names the best system prompt (flagging close calls) and returns its reusable prompt_id; failed variants are named, not given invented scores.

Where this leads

Natural next recipes once you are comfortable with this one.

Recipes that lead here