Changelog
What’s new
2026-09
- Benchmark cost estimator – see estimated cost before running a suite, including judge and diagnostic scoring
- Refresh buttons – one-click refresh on Comparisons, Evaluations, and Benchmark run history
- Model substitution – deprecated models are automatically substituted at run time; results show what was used
- BYOE endpoints – bring your own OpenAI-compatible model endpoint and benchmark it alongside hosted providers
2026-08
- Initial launch – multi-model comparison, rubric evaluation, benchmark suites, RAG context