We ran a benchmark. The scores were wrong (but the models were fine).
Seven models, three runs each. The task: summarise a customer support ticket into three bullet points. The prompt and ticket are shown below.
Scores came back ranging from 0 to 83.3. GPT-4.1 Nano scored 83.3 on run 2 and 0 on runs 1 and 3. Claude Opus 5 scored 83.3, then 62.5, then 45.8, declining across three consecutive runs on a simple task.