MCP Recipes
ProOptimize a Prompt in a Loop
Iteratively improve a prompt: A/B the current best against an agent-proposed variant, keep the winner, repeat -- stopping at a quality target, a plateau, an iteration cap, or a cost cap. A closed optimization loop with the judge as the fitness function.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, run a closed-loop optimization of a prompt. You are the optimizer; the scoring (gold standard or rubric) is the fitness function. This recipe COMPOSES scored evaluations with an agent-state note that holds the loop's state so it can run across sessions. Keep status "running" until the loop ends. SET UP: 1. Confirm the tools you need are available (list the tools). 2. Read your notes: get_agent_notes(recipe_id="recipe_id"). If an active note exists, we are RESUMING a loop -- read its data (iteration, best, budget, history) and continue from "LOOP". If none, we are STARTING -- continue here. 3. Ask me for: the starting prompt; the ONE model to optimize for; the scoring standard (expected_output gold answer, OR a saved rubric_id); and the STOP CONDITIONS -- a target score, a max iteration count, and a cost cap (dollars). Optionally a plateau threshold (minimum score improvement worth continuing). If I do not give budgets, propose sensible defaults and state them explicitly -- NEVER run an unbounded loop. 4. Record the loop state in a note BEFORE spending: put_agent_note(recipe_id="recipe_id", recipe="mcp-prompt-optimizer", status="running", intent="Optimize prompt for <model> toward <target>", data={ model, scoring:{expected_output|rubric_id}, target, max_iterations, cost_cap, plateau, iteration:0, cost_spent:0, best:{prompt:<starting prompt>, score:null, job_id:null}, history:[] }). LOOP (each iteration): 5. SCORE the current best (if not yet scored this iteration): run_evaluation(prompt=<best prompt>, models=<the model>, expected_output|rubric_id=<scoring>). Async -- poll_job every 5 seconds to completion. Read EXACTLY results[].quality_score for the model (0-100). 6. CHECK STOP CONDITIONS against the note's budget -- stop and go to REPORT if ANY is true: - best.score >= target, OR - iteration >= max_iterations, OR - cost_spent >= cost_cap, OR - (if plateau set and iteration > 0) the score improved by less than plateau since the last iteration. Do the comparison explicitly on the numbers; do not eyeball it. 7. GENERATE one improved variant: propose a single, specific change to the best prompt, INFORMED by the judge's reasoning on what the current best lost points for (read the breakdown/judge_reasoning). State the hypothesis ("adding an explicit format instruction should lift the structure score"). Change ONE thing -- this keeps it a clean A/B. 8. A/B: score the new variant the same way (run_evaluation, same model + scoring). Read its quality_score the same way (results[].quality_score). 9. SELECT the winner: the higher quality_score of {current best, new variant}. If within the plateau threshold, treat as no improvement. 10. UPDATE the note (this is the loop's memory -- get it right): set iteration += 1; cost_spent += this iteration's cost (sum the run costs); best = the winner {prompt, score, job_id}; append {iteration, variant, score, cost} to history. Then READ THE NOTE BACK (get_agent_notes) and confirm the written state matches what you intended -- if it does not, correct it before continuing. Do NOT carry loop state only in your own working memory; the note is the source of truth. 11. Go back to step 6 (re-check stop conditions, then iterate). REPORT (when a stop condition fired): 12. Tell me: the best prompt, its score, WHICH stop condition ended the loop, the full iteration history (score per iteration so I can see the trajectory), and total cost_spent. 13. OVERFIT CHECK (important): a high judge score is not proof the prompt is genuinely better -- an optimizer can learn to please the judge (verbose, judge-flattering output) rather than improve real quality. Show me the actual best response, not just the number, and recommend I eyeball it before adopting. Do NOT present the optimized prompt as definitively best on score alone. 14. CLEANUP: mark the note "completed" (it drops from the active list, frees a slot). If I want to keep optimizing later, leave it "running" as the series anchor instead and tell me so. ON ANY FAILURE or clear mismatch (a run will not score, the note write did not read back correctly, you lost track of iteration state): STOP the loop, mark the note "failed", and stamp data with failure_stage, failure_reason, a one-line detail, and the last good job_id. Never keep spending in a loop whose state you cannot trust, and never report an optimized result you did not actually measure.
Goal
Improve a prompt by iterating, not guessing. The agent A/Bs the current best against a proposed improvement, keeps the winner, and repeats – converging toward a quality target within budgets you set. It is the agent acting as an optimizer with the scoring as its fitness function, and a durable note as its memory so the loop can run across sessions.
When to use
Reach for this when a prompt matters enough to tune properly and you want evidence-driven iteration rather than one-shot rewriting. Best when you have a clear scoring standard (a gold answer or a rubric) so “better” is a number, not an opinion.
How it works, and its limits (read this)
- The loop lives in prose + a note. There is no special loop engine: the recipe steps are the control flow, the agent is the executor, and the agent-state note holds iteration / best / cost / history so a fresh session resumes exactly where it left off. The note is the source of truth – the agent writes it and reads it back each iteration to stay honest.
- Hard stops are mandatory. The loop ends on the FIRST of: target reached, plateau (gains too small to bother), iteration cap, or cost cap. There is no unbounded run. Budgets live in the note so they survive sessions.
- The judge is an imperfect fitness function. Optimizing to a score can produce prompts that please the judge without being genuinely better (Goodhart’s law). The recipe ends with a human checkpoint: look at the actual best output, not just the number, before adopting.
- Cost is tracked and reported. Every iteration is scored runs that cost money; the loop reports cumulative spend against your cap and stops when it is hit.
Verification
- The loop had explicit stop conditions (target / iterations / cost, optionally plateau) set before any spend, held in the note.
- Each iteration updated the note and read it back; loop state did not live only in working memory.
- The final report names the best prompt, its score, which stop condition fired, the per-iteration trajectory, and total cost.
- A human checkpoint on the actual output (not just the score) was offered before adoption.
- On a state or scoring failure, the loop stopped and the note was marked failed – it did not keep spending.
Where this leads
Natural next recipes once you are comfortable with this one.