You don't need to generate to evaluate
Every evaluation tool starts by running a model. You give it a prompt, it generates an answer, and then it scores what came out. That is the right shape when the thing you want to judge does not exist yet. It is the wrong shape when you already wrote the thing.
Most of the time, you have the text. A tagline is sitting in a doc. A support reply is drafted and waiting to send. A clause is written and you need to know if it holds up. You do not want a model to generate anything. You want a judgment on the words you already have.
Scoring is a new run type that does exactly that. It judges supplied text against a rubric or your compliance documents, and it runs no model. There is no generation step, so there is no cost, no latency, and no model in the result. It scores what you paste, returns a score per criterion, and calls a verdict. It sits alongside comparison, evaluation, and benchmarking as a first-class run type, with its own tile, its own tool, and its own result view.
“Is this compliant?”
Here is the use case that proves the point, run for real against our own marketing compliance standard. Start with a line of product copy that looks completely fine at a glance.
The copy under review
Read that as a human and it reads like confident marketing. Score it against a compliance rubric backed by our regulation documents, and the picture sharpens. The rubric checks four clauses: no unbacked superlatives, feature availability stated per tier, accurate pricing, and no fabricated claims.
The verdict (per clause, with the regulation each was judged against)
- Superlatives: fail. “most accurate” and “always” are absolute performance claims with no cited benchmark data (Clause 1.3).
- Tier clarity: fail. RAG context and compliance scoring are gated features, and the copy names no tier (Clause 1.1).
- Pricing: pass. No pricing figures stated.
- Fabrication: pass. No testimonials or invented claims.
The copy scored 50 out of 100. Two flags, each tied to a specific clause, each quoting the exact phrase at fault. This is the output a marketer actually needs: not “looks fine” but “here are the two things that will not survive review, and here is the rule each one breaks.”
The fix follows directly from the flags. Drop the superlatives, state the tier.
The revised copy
Scored again, the revised line comes back at 100 out of 100, all four clauses clear. Same rubric, same regulation documents, a one-line edit between them.
Compliance score, same copy, before and after one fix (out of 100)
Per-clause scores for the same copy, before and after a single edit
“Which phrasing is best?”
The second use case: stack-rank several versions of the same thing. Paste two taglines, three clause drafts, four support replies, and each one is scored independently against the same standard, then a winner is called. Here are four taglines we scored against a brand-voice rubric (lead with the reader’s outcome, be concrete, avoid hype, read cleanly):
The four taglines scored
- “LLM Prover Score: rubric-based and compliance scoring for text artifacts, with per-criterion results.” (product-led, feature list)
- “Judge the copy you already wrote, against your own rubric, in seconds.” (benefit-led, plain)
- “Score the copy you already have against your own rubric, and see exactly which lines fall short.” (benefit-led, more detail)
- “The most powerful way ever to check if your writing is good enough to ship.” (hype)
Four taglines, scored against one brand-voice rubric (out of 100)
Four taglines, one rubric, scored independently
The ranking is useful, but the more useful thing is what the per-criterion detail showed. The benefit-led line “Score the copy you already have against your own rubric, and see exactly which lines fall short” scored lower than expected. The breakdown explained why: the rubric’s “concrete and specific” criterion had been written in a way that rewarded naming product features and quietly penalised a line that led with the reader’s benefit instead. The rubric was pulling against itself, rewarding one thing in the “outcome” criterion and the opposite in the “concrete” criterion.
That is a problem with the standard, not the copy. The rubric was rewritten so that “concrete” means vivid and specific language, whether it describes the reader’s action or the product. Scored again, the ranking settled into something sensible, and the benefit-led line rose to match.
Score the copy you already have
Scoring is on Pro and above. Compliance scoring against your own regulation documents is on Pro Plus and above.
No model, no cost, nothing faked
The result view for a scoring run shows no cost, no latency, and no model. That is deliberate. Nothing was generated, so there is nothing to price or time.
What decides the verdict is your standard. Provide a quality rubric and the judge scores your text against your criteria. Provide a compliance store and it judges policy conformance against your regulation documents. Same text, different standard, and the standard is yours to define.
One honest boundary on the compliance use case: a compliance score is an indicator, not a certification. It tells you where copy conflicts with the clauses you supplied, which is a strong signal for a human reviewer, not a stamp of legal approval. Read the per-clause reasoning, that is where the value is.
Where to reach it
Scoring, like comparisons and evaluations, works anywhere you do:
- Dashboard: a new scoring tile. Paste one artifact, or several to stack-rank, pick a rubric or compliance store, and run.
- Agent and MCP:
scoreis a core tool, so “check if this copy is compliant” or “score these three taglines against our brand rubric” works directly from your agent in Slack or your IDE. - API:
POST /score, the same async pattern as the other run types.
Scoring is available from Pro and up, with compliance scoring on Pro Plus.
The fastest path from “I wrote this” to “is it good, is it compliant, which version wins” no longer runs through a model. You already have the text. Now you can judge it.
What’s next
Evaluations Guide
Score model outputs against a rubric or gold standard answer and get per-criterion reasoning.
Compliance Store Guide
Upload regulations or policies as a compliance store and score model outputs against specific clauses.
Rubric Guide
How to write criteria that produce reliable, consistent judge scores, and how to fix them when they don't.