The $323 model and the $0.60 model scored the same. Until the task got harder.
We ran a root cause classification task across seven models, four times each. The task was designed to be hard in a specific way: five support tickets, all describing problems with the same product feature, but with different root causes buried in timezone and timing signals. A human reads it and sees the pattern immediately. An LLM has to reason across all five tickets simultaneously, convert local check times to UTC, and resist grouping by surface symptom rather than mechanism.