MCP Recipes
ProA/B Test Your System Prompt
Save two or more system prompts, run the same task through each on one model, score them the same way, and see which set of instructions produces the best result.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, run a single-variable A/B test on the SYSTEM PROMPT: same model, same task prompt, same scoring, only the system prompt changes. 1. Confirm the tools you need are available (list the tools). 2. Ask me for the system prompt variants to test (two or more) and the task prompt they will all run against, if I have not given them. 3. Lock the controls so this is a fair test: pick ONE model (list_models; sensible default or ask me) and ONE scoring method -- a gold-standard answer (expected_output) or a saved rubric (rubric_id). Model, task prompt, and scoring stay identical across variants. 4. Save each system prompt variant: create_system_prompt(name, content) -- note each prompt_id. (Saving them makes the winner reusable afterward, and keeps the test clean.) Respect my plan's system-prompts cap; if I am near it, say so. 5. Run each variant as its own scored evaluation, identical except for the system prompt: for each, run_evaluation(prompt=<the task prompt>, models=<the one model>, system_prompt_id=<that variant's id>, expected_output=<gold> OR rubric_id=<rubric>). Each is async -- poll_job every 5 seconds to completion (stop after 5 minutes). 6. Compare scores across variants and recommend the best system prompt. Show each variant's quality_score side by side with the judge's reasoning and name the winner; if two are within a few points, call it close rather than overstating a small difference. 7. Hand back the winning system prompt's prompt_id and remind me it is saved and reusable (attach it via system_prompt_id in any comparison, evaluation, or benchmark). Offer once to monitor the winning setup over time, if it fits. ON ANY FAILURE or clear mismatch (a create fails, a variant will not score): tell me honestly what happened and at which step -- compare the variants that succeeded, name the ones excluded, and do not invent a score. If some system prompts were saved before a failure, tell me their prompt_ids so we do not create duplicates on retry.
Goal
Find the system prompt that gets the best behaviour out of a model for your task. You give several sets of instructions; the agent saves each, runs the same task under each on the same model, scores them the same way, and tells you which system prompt wins – with the winner left saved and reusable.
When to use
Reach for this when the model’s behaviour depends on how you instruct it and you want to choose the best instructions with evidence. It is the system-prompt-stage A/B test: fix the model, the task, and the scoring; vary only the system prompt.
How it works
- One variable. Model, task prompt, and scoring are fixed; only the system prompt changes, so score differences are attributable to the instructions.
- Variants are saved, not throwaway. Each system prompt is saved via create_system_prompt, so the winner is immediately reusable (and the test references stable prompt_ids).
- Scored and honest. Each variant gets a quality_score and reasoning; close results are called close, not spun into a winner.
Verification
- Every variant used the same model, task prompt, and scoring – only the system prompt differed.
- Each system prompt was saved (has a prompt_id) and has its own score.
- The recommendation names the best system prompt (flagging close calls) and returns its reusable prompt_id; failed variants are named, not given invented scores.
Where this leads
Natural next recipes once you are comfortable with this one.
Lock in the winning system prompt and monitor the task over time, so you catch any quality drift.
Judge the variants on exactly what matters -- build a reusable rubric and re-run the A/B against it.
Recipes that lead here