Watch an agent tune a system prompt, one scored iteration at a time
A support-triage system prompt scored zero on the first run. Two iterations later it scored a perfect hundred. The agent made both changes on its own, and each one was informed by exactly why the previous version lost points. No blind rewrites. A loop that reasons from evidence, stops when it hits a target, and shows every step.
This is the “Optimize a Prompt in a Loop” recipe. The agent is the optimizer. A scored evaluation is the fitness function. The run below is real, start to finish.