Benchmarks
Benchmark Guide
Save a run configuration, schedule it on a cadence, and track quality and cost drift over time.
This guide is coming soon. In the meantime, see the Getting Started guide for an overview of how benchmarks work.
Start tracking model drift
Pro and above. Schedule benchmarks and get alerted the moment quality drops.
See pricingWhat’s next
Getting Started
The three ways to run a prompt, the config items that sharpen results, and catch AI model drift before users do.
Understanding Drift
What model drift is, how LLM Prover detects it, and how to act on anomaly flags.
Scoring Guide
How LLM Prover calculates scores, what affects them, and how to use the diagnostic judge to find out why.
Evaluations Guide
Score model outputs against a rubric or gold standard answer and get per-criterion reasoning.