Skip to content

MCP Recipes

Pro

A/B Test Your Prompt

Try several wordings of the same request against one model, score them the same way, and see which phrasing produces the best result -- a single-variable A/B test on your prompt.

Run it: paste into your MCP agent

Highlighted parts are placeholders. Replace them with your own values before (or after) copying.

Using the LLM Prover MCP tools, run a single-variable A/B test on PROMPT WORDING: same
model, same scoring, only the phrasing changes. The point is to isolate the effect of how
the request is worded.

1. Confirm the tools you need are available (list the tools).
2. Ask me for the prompt wordings to test (two or more variants of the same request) if I
   have not given them. These are the only thing that varies.
3. Lock the controls so this is a fair test: pick ONE model (list_models; choose a sensible
   default or ask me) and ONE scoring method -- a gold-standard answer (expected_output) or
   a saved rubric (rubric_id). The same model and the same scoring apply to every variant.
4. Run each wording as its own scored evaluation, keeping model and scoring identical:
   for each variant, run_evaluation(prompt=<that wording>, models=<the one model>,
   expected_output=<gold> OR rubric_id=<rubric>). Each is async -- poll_job every 5 seconds
   to completion (stop after 5 minutes). Run them one per variant so each gets its own score.
5. Compare the scores across wordings and recommend the best-performing phrasing. Show me
   each variant's quality_score side by side with the judge's reasoning, and name the
   winner. If two are within a few points, say so -- call it close rather than overstating
   a tiny difference, and note cost/latency if they differ meaningfully.
6. Offer the next step once: A/B the system prompt the same way, or monitor the winning
   prompt over time -- only if it fits.

ON ANY FAILURE or clear mismatch (a variant will not score, the model errors on one): tell
me honestly which variants scored and which did not -- compare the ones that succeeded and
name the ones excluded, do not invent a score for a failed run.

Goal

Find the best way to word a request by testing it. You give several phrasings of the same ask; the agent scores each against the same model and the same standard and tells you which wording wins – isolating phrasing as the single variable so the comparison is fair.

When to use

Reach for this when output quality seems sensitive to how you phrase the request and you want evidence, not a hunch. It is the prompt-stage sibling of A/B-testing a system prompt or comparing models: hold everything else constant, vary one thing, measure.

How it works

  • One variable. The model and the scoring are fixed; only the prompt wording changes, so any score difference is attributable to the phrasing.
  • Scored, not eyeballed. Each wording runs through a scored evaluation (gold standard or rubric), so the comparison is a number plus the judge’s reasoning, not a vibe.
  • Honest about close calls. A two-point gap is not a clear winner; the recipe says so rather than overstating noise.

Verification

  • Every variant used the same model and the same scoring method – only the prompt differed.
  • Each wording has its own quality_score and the recommendation names the best, flagging close results.
  • Failed variants are named and excluded, not given invented scores.

Where this leads

Natural next recipes once you are comfortable with this one.