Benchmarks
Understanding Drift
What model drift is, how LLM Prover detects it, and how to act on anomaly flags.
This guide is coming soon. In the meantime, see the Benchmark Guide for how to set up scheduled runs.
Catch drift before users do
Pro and above. Scheduled benchmarks alert you the moment a model update changes your results.
See pricingWhat to do next
Benchmark Guide
Save a run configuration, schedule it on a cadence, and track quality and cost drift over time.
Scoring Guide
How LLM Prover calculates scores, what affects them, and how to use the diagnostic judge to find out why.
Evaluations Guide
Score model outputs against a rubric or gold standard answer and get per-criterion reasoning.
Getting Started
An orientation to LLM Prover -- run types, config, and how the pieces fit together.