Evaluation
Evaluations Guide
Score model outputs against a rubric or gold standard answer and get per-criterion reasoning.
This guide is coming soon. In the meantime, see the Getting Started guide for an overview of how evaluations work, and the Scoring Guide for how scores are calculated.
What to do next
Getting Started
An orientation to LLM Prover -- run types, config, and how the pieces fit together.
Scoring Guide
How LLM Prover calculates scores, what affects them, and how to use the diagnostic judge to find out why.
Rubric Writing Guide
How to write criteria that produce reliable, consistent judge scores -- and how to fix them when they don't.
Comparisons Guide
Run a prompt against multiple models simultaneously and compare cost, latency, and output side by side.