Evaluation
Evaluations Guide
Score model outputs against a rubric or gold standard answer and get per-criterion reasoning.
This guide is coming soon. In the meantime, see the Getting Started guide for an overview of how evaluations work, and the Scoring Guide for how scores are calculated.
What’s next
Getting Started
The three ways to run a prompt, the config items that sharpen results, and catch AI model drift before users do.
Scoring Guide
How LLM Prover calculates scores, what affects them, and how to use the diagnostic judge to find out why.
Rubric Writing Guide
How to write criteria that produce reliable, consistent judge scores -- and how to fix them when they don't.
Comparisons Guide
Run a prompt against multiple models simultaneously and compare cost, latency, and output side by side.