LLM Prover is now programmable
LLM Prover started as a dashboard. You opened it, ran a comparison, read the results. That workflow works. It also has a ceiling: it requires a human to initiate every run, read every result, and decide what to do next.
The REST API removes that ceiling. Every operation available in the dashboard is now available as an API call. Your scripts, pipelines, and CI workflows can fire comparisons, run evaluations, trigger benchmarks, and retrieve results without a human in the loop.
What shipped
A full REST API covering every core operation:
POST /compare– fire a prompt against multiple models in parallelPOST /evaluate– score a response against a rubricPOST /benchmarks/{id}/run– trigger a benchmark suite runGET /jobs/{job_id}– poll an async jobGET /compare/{id},GET /evaluations/{id},GET /benchmarks/{id}/runs– retrieve resultsGET /models– list available models with current pricingGET /usage– retrieve usage and cost data programmatically
Developer keys are available on all plans. Starter gets one key. Pro gets five. Enterprise gets fifty. Keys are managed in your dashboard settings.
Full reference at the API docs.
The async pattern
Every API call that starts a run returns a job_id immediately. You poll GET /jobs/{job_id} until status is complete, then retrieve the full result.
Start a comparison:
POST /compare with your prompt and models list – returns { "job_id": "..." }
Poll:
GET /jobs/{job_id} – returns { "status": "pending" | "complete" | "failed" }
Retrieve:
GET /compare/{comparison_id} – returns full result with response text, scores, cost, latency
This pattern is the same whether you are running a single comparison or a full benchmark suite with twenty prompts. It is also the same pattern the MCP server uses internally, which means scripts and agents interoperate with the same result format.
What you can build with it
The use cases split across three levels.
Individual developers and CI pipelines
Quality gates on every PR. A GitHub Action calls POST /evaluate on every pull request. If the quality score drops below your threshold, the job fails and the merge is blocked. The rubric lives in your repo alongside your code. This applies to model changes and prompt changes equally – if prompts live in your repo (they should), every PR that touches a prompt file triggers an evaluation. The prompt library has the same quality guarantees as the code it ships with.
Canary deploys for model upgrades. Before cutting over to a new model, run both the old and new model against your benchmark suite via the API. Only switch when the new model meets or exceeds the old one’s scores. GET /models returns current pricing; POST /compare returns per-model cost and latency. The upgrade decision is scripted and repeatable, not a manual dashboard exercise.
Production monitoring. A cron job calls POST /benchmarks/{id}/run nightly. Alert thresholds fire the moment a run breaches your quality floor. You find out about model drift before your users do.
Teams
Bulk evaluation pipelines. A data team scripts evaluations across hundreds of prompts, exports results to their BI stack via GET /usage, and tracks quality trends over time. The same rubric runs on every item. No reviewer fatigue, no sampling.
Dynamic model routing. Before dispatching a task, your pipeline calls POST /compare on the actual input across candidate models and routes to the winner. The routing decision is based on live data from the real task, not a static config. As models update and costs shift, the routing adapts.
Inline output scoring. Your application generates a response, immediately calls POST /evaluate, and uses the score to decide what to do next – surface the response, retry with a different model, or escalate to a human. LLM Prover becomes a quality gate inside your request path.
Platforms
Embedded measurement. A B2B SaaS company building on top of LLMs embeds LLM Prover as their eval infrastructure. Their customers get quality scores on their own data. The SaaS company doesn’t build eval from scratch – they script it per customer via the API. Every comparison and evaluation result is retrievable by ID, which means every model decision has an audit trail. For regulated industries, that is a compliance argument: “we evaluated this output against this rubric on this date and it scored X.”
Compliance verification pipelines. A document processing pipeline calls POST /evaluate with a compliance rubric before any generated document leaves the system. Documents that fail are held for human review. The pipeline never dispatches a non-compliant document downstream.
If your client is an agent
The REST API requires integration code. You write the HTTP calls, handle the polling loop, parse the response. That is the right choice for pipelines and scripts where you control the execution environment.
If your client is an AI agent (Claude Desktop, Cursor, VS Code Copilot, Amazon Q), the MCP server is the faster path. The agent discovers the 22 available tools automatically and calls them in natural language. No integration code. The same async pattern runs under the hood.
Both paths return the same results. The difference is who writes the integration: you, or the agent.
See the MCP launch post for a live example of an agent running a comparison end to end.
Add LLM Prover to your pipeline
REST API on all plans. Developer key in your dashboard settings.
Get started