User Guides
Evaluation
Scoring Guide
How LLM Prover calculates scores, what affects them, and how to use the diagnostic judge to find out why.
Comparisons Guide
Run a prompt against multiple models simultaneously and compare cost, latency, and output side by side.
Diagnostics Guide
How the diagnostic judge works, what it flags, and how to act on findings in comparisons, evaluations, and benchmarks.
Evaluations Guide
Score model outputs against a rubric or gold standard answer and get per-criterion reasoning.
Prompting Guide
How to write prompts that produce reliable, comparable results across models.
Config
System Prompts Guide
When to use system prompts, how to write them, how they interact with rubrics and context stores, and how to A/B test two system prompts on the same model.
Rubric Writing Guide
How to write criteria that produce reliable, consistent judge scores -- and how to fix them when they don't.
RAG Store Guide
How to create a context store, upload documents, and get reliable retrieval in comparisons, evaluations, and benchmarks.
Compliance Store Guide
Upload regulations or policies as a compliance store and score model outputs against specific clauses.
BYOE Guide
Register any OpenAI-compatible endpoint and benchmark your own models against any model in the registry.
Benchmarks
Benchmark Guide
Save a run configuration, schedule it on a cadence, and track quality and cost drift over time.
Alerts Guide
How to configure alert thresholds, what each alert type means, and how to act on score_drop, cost_spike, latency_spike, and model_version_change alerts.
Understanding Drift
What model drift is, how LLM Prover detects it, and how to act on anomaly flags.
No docs match your search.