Config
BYOE Guide
Register any OpenAI-compatible endpoint and benchmark your own models against any model in the registry.
This guide is coming soon. In the meantime, see the Getting Started guide for an overview of how endpoints work.
Benchmark your own models
Pro Plus and above. Register your endpoint and prove your fine-tune is better.
See pricingWhat to do next
Getting Started
An orientation to LLM Prover -- run types, config, and how the pieces fit together.
Benchmark Guide
Save a run configuration, schedule it on a cadence, and track quality and cost drift over time.
Evaluations Guide
Score model outputs against a rubric or gold standard answer and get per-criterion reasoning.
Compliance Store Guide
Upload regulations or policies as a compliance store and score model outputs against specific clauses.