Skip to content

Watch an agent tune a system prompt, one scored iteration at a time

· 7 min read

A support-triage system prompt scored zero on the first run. Two iterations later it scored a perfect hundred. The agent made both changes on its own, and each one was informed by exactly why the previous version lost points. No blind rewrites. A loop that reasons from evidence, stops when it hits a target, and shows every step.

This is the “Optimize a Prompt in a Loop” recipe. The agent is the optimizer. A scored evaluation is the fitness function. The run below is real, start to finish.

The job: triage tickets into clean JSON

A ticket router needs structured input. Given a customer message, a classifier has to return a single JSON object with three fields: a category, a severity, and a one-line summary. A downstream system reads that JSON and routes the ticket. Prose breaks it. A wrong severity sends an outage to the bottom of the queue.

The test message is deliberately messy, and severity needs reasoning rather than keyword matching.

The ticket

hi, so i was trying to export my quarterly report last night and the whole dashboard just froze, then logged me out. tried again this morning, same thing, cant get in at all now. i have a board meeting at 2pm and that report is the only thing on the agenda. please help

A rubric defines “good” so the loop has a number to optimize against. Four criteria: valid JSON structure, correct category from a fixed set, correct severity by business impact, and a concise accurate summary. That rubric is the fitness function for the whole run.

Iteration 0: the starting point

The first system prompt is deliberately thin.

System prompt: You are a support assistant. Classify the user's ticket. 0
Category: Technical Issue - Dashboard Freezing and Login Problem

Zero. The model returned a label, not JSON. No severity, no summary, nothing a router can parse. The judge did not just score it down. It pointed at the cause: the system prompt never asked for a JSON object, never listed the keys, never named the allowed values.

That diagnosis is the input to the next move.

The agent changes one thing, and the score moves

Iteration one adds the missing contract: return a single JSON object, these exact keys, no prose, no code fences. One change, so the result stays readable.

Iteration 1: + JSON contract 55
{"category":"Technical Issue","severity":"High","summary":"Dashboard freezes and logs out user when attempting to export quarterly report; user unable to access dashboard, urgent due to upcoming board meeting."}

Fifty-five, up from zero. The structure is right and the severity reasoning is right: the judge noted it correctly assigned high severity to a service outage blocking critical business access. Two criteria still failed, and the judge was specific about both. The category read Technical Issue, not the required enum value technical. The summary ran well past the fifteen-word limit.

So iteration two makes one more targeted change: list the allowed category values, and constrain the summary to a single short sentence.

Agent
> run_evaluation (best prompt, rubric)
score: 55 | category fail, summary fail
judge docked category (off-enum) and summary (too long). Change one thing: add enums + length limit.
> create_system_prompt (variant)
> run_evaluation (variant, rubric)
score: 100 | all criteria clean
variant beats best. Keep it. Target reached, stop.
> update_agent_note (iteration, best, history, cost)
Iteration 2: + enums and summary limit 100
{"category":"technical","severity":"high","summary":"Dashboard freezes and logs out user, blocking access to quarterly report."}

A hundred. Every criterion clean. The judge flagged the category as a correct match: technical fits the described dashboard freeze and access failure, rather than a surface keyword like “logout.” The model classified on intent, not vocabulary, and severity stayed high because a total lockout against a hard deadline is a real business impact.

The whole trajectory, driven entirely by the judge’s per-criterion feedback:

IterationChange to the system promptScoreJSONCategorySeveritySummary
0classify the ticket0failfailfailfail
1require a JSON object, fixed keys55passfailpassfail
2add category enums, limit the summary100passpasspasspass

Each change targeted the exact criterion the judge had just docked

Why this is a loop, not a macro

Nothing new was built to make this run. It composes three things that already exist: a scored evaluation, a durable note that holds the loop’s state, and an agent that repeats until a stop condition. The agent proposes a change, scores it, keeps the better prompt, and reads the judge’s reasoning to decide what to try next. That is the loop.

The note is what makes it durable. The recipe instructs the agent to write iteration, current best, history, and spend into a note after every step, then read it back to stay honest. Start a loop, close your laptop, resume in a fresh session, and the agent knows exactly where it was. The note is the memory, not the agent’s short-term context.

Every loop is bounded before it spends a cent. It stops on the first of four conditions: a target score, a plateau where gains are too small to bother, an iteration cap, or a cost cap. There is no unbounded run and no surprise bill. In this case the target fired at iteration two.

Run the optimization loop on your own prompt

The Optimize a Prompt in a Loop recipe is on Pro and above. Bring a prompt and a scoring standard; your agent does the iterating.

Get started on Pro

The score is not the point. The output is.

The loop climbs a number, and a number is a proxy. Any optimizer pointed at a proxy will eventually learn to satisfy the proxy rather than the goal. A prompt can learn to flatter the judge, scoring well on verbosity or structure while drifting away from what the output actually needs to do. The score keeps rising. The result stops improving.

This is why the recipe ends with a human checkpoint. The loop does the iterating and reports the trajectory, but the final call on whether the winning prompt is genuinely better belongs to a person reading the actual output. The score narrows the field. It does not close the decision.

Insight: Optimizing to a score rewards whatever the score measures, including its blind spots. A high number is a reason to look at the output, not a substitute for looking at it.

What it cost, and what you keep

The converging run cost a fraction of a cent. At the end you have two things: a system prompt that scores a measured hundred instead of a guessed-at zero, and the full trail of how it got there, iteration by iteration, with the judge’s reasoning attached to each step.

The one prerequisite is a scoring standard. The loop needs a gold-standard answer or a rubric, because “better” has to be a number before an agent can climb toward it. Give it that, and the agent does the iterating that you would otherwise do by hand.

This run used a judge as the fitness function. The next step in the story is a loop that optimizes against ground truth instead of an opinion: a code sandbox where the agent tunes a function until the tests pass and it runs fast enough. Judge versus reality. That is a post for another day.

What’s next

mcp evaluation agentic product-update