Evaluation
Comparisons Guide
Run a prompt against multiple models simultaneously and compare cost, latency, and output side by side.
This guide is coming soon. In the meantime, see the Getting Started guide for an overview of how comparisons work.
What to do next
Getting Started
An orientation to LLM Prover -- run types, config, and how the pieces fit together.
Scoring Guide
How LLM Prover calculates scores, what affects them, and how to use the diagnostic judge to find out why.
Evaluations Guide
Score model outputs against a rubric or gold standard answer and get per-criterion reasoning.
Benchmark Guide
Save a run configuration, schedule it on a cadence, and track quality and cost drift over time.