Watch an agent tune a system prompt, one scored iteration at a time
A support-triage system prompt scored zero on the first run. Two iterations later it scored a perfect hundred. The agent made both changes on its own, and each one was informed by exactly why the previous version lost points. No blind rewrites. A loop that reasons from evidence, stops when it hits a target, and shows every step.
This is the “Optimize a Prompt in a Loop” recipe. The agent is the optimizer. A scored evaluation is the fitness function. The run below is real, start to finish.
The job: triage tickets into clean JSON
A ticket router needs structured input. Given a customer message, a classifier has to return a single JSON object with three fields: a category, a severity, and a one-line summary. A downstream system reads that JSON and routes the ticket. Prose breaks it. A wrong severity sends an outage to the bottom of the queue.
The test message is deliberately messy, and severity needs reasoning rather than keyword matching.
The ticket
A rubric defines “good” so the loop has a number to optimize against. Four criteria: valid JSON structure, correct category from a fixed set, correct severity by business impact, and a concise accurate summary. That rubric is the fitness function for the whole run.
Iteration 0: the starting point
The first system prompt is deliberately thin.
Zero. The model returned a label, not JSON. No severity, no summary, nothing a router can parse. The judge did not just score it down. It pointed at the cause: the system prompt never asked for a JSON object, never listed the keys, never named the allowed values.
That diagnosis is the input to the next move.
The agent changes one thing, and the score moves
Iteration one adds the missing contract: return a single JSON object, these exact keys, no prose, no code fences. One change, so the result stays readable.
{"category":"Technical Issue","severity":"High","summary":"Dashboard freezes and logs out user when attempting to export quarterly report; user unable to access dashboard, urgent due to upcoming board meeting."}Fifty-five, up from zero. The structure is right and the severity reasoning is right: the judge noted it correctly assigned high severity to a service outage blocking critical business access. Two criteria still failed, and the judge was specific about both. The category read Technical Issue, not the required enum value technical. The summary ran well past the fifteen-word limit.
So iteration two makes one more targeted change: list the allowed category values, and constrain the summary to a single short sentence.
{"category":"technical","severity":"high","summary":"Dashboard freezes and logs out user, blocking access to quarterly report."}A hundred. Every criterion clean. The judge flagged the category as a correct match: technical fits the described dashboard freeze and access failure, rather than a surface keyword like “logout.” The model classified on intent, not vocabulary, and severity stayed high because a total lockout against a hard deadline is a real business impact.
The whole trajectory, driven entirely by the judge’s per-criterion feedback:
| Iteration | Change to the system prompt | Score | JSON | Category | Severity | Summary |
|---|---|---|---|---|---|---|
| 0 | classify the ticket | 0 | fail | fail | fail | fail |
| 1 | require a JSON object, fixed keys | 55 | pass | fail | pass | fail |
| 2 | add category enums, limit the summary | 100 | pass | pass | pass | pass |
Each change targeted the exact criterion the judge had just docked
Why this is a loop, not a macro
Nothing new was built to make this run. It composes three things that already exist: a scored evaluation, a durable note that holds the loop’s state, and an agent that repeats until a stop condition. The agent proposes a change, scores it, keeps the better prompt, and reads the judge’s reasoning to decide what to try next. That is the loop.
The note is what makes it durable. The recipe instructs the agent to write iteration, current best, history, and spend into a note after every step, then read it back to stay honest. Start a loop, close your laptop, resume in a fresh session, and the agent knows exactly where it was. The note is the memory, not the agent’s short-term context.
Every loop is bounded before it spends a cent. It stops on the first of four conditions: a target score, a plateau where gains are too small to bother, an iteration cap, or a cost cap. There is no unbounded run and no surprise bill. In this case the target fired at iteration two.
Run the optimization loop on your own prompt
The Optimize a Prompt in a Loop recipe is on Pro and above. Bring a prompt and a scoring standard; your agent does the iterating.
The score is not the point. The output is.
The loop climbs a number, and a number is a proxy. Any optimizer pointed at a proxy will eventually learn to satisfy the proxy rather than the goal. A prompt can learn to flatter the judge, scoring well on verbosity or structure while drifting away from what the output actually needs to do. The score keeps rising. The result stops improving.
This is why the recipe ends with a human checkpoint. The loop does the iterating and reports the trajectory, but the final call on whether the winning prompt is genuinely better belongs to a person reading the actual output. The score narrows the field. It does not close the decision.
What it cost, and what you keep
The converging run cost a fraction of a cent. At the end you have two things: a system prompt that scores a measured hundred instead of a guessed-at zero, and the full trail of how it got there, iteration by iteration, with the judge’s reasoning attached to each step.
The one prerequisite is a scoring standard. The loop needs a gold-standard answer or a rubric, because “better” has to be a number before an agent can climb toward it. Give it that, and the agent does the iterating that you would otherwise do by hand.
This run used a judge as the fitness function. The next step in the story is a loop that optimizes against ground truth instead of an opinion: a code sandbox where the agent tunes a function until the tests pass and it runs fast enough. Judge versus reality. That is a post for another day.
What’s next
MCP Recipes Guide
Hand your AI agent a ready-to-paste playbook that runs LLM Prover tools in order, polls for results, and reports back in plain language.
Rubric Guide
How to write criteria that produce reliable, consistent judge scores, and how to fix them when they don't.
Monitoring Guide
How LLM Prover detects changes in model quality, cost, and latency, and how to configure alerts so you know before users do.