Skip to content

You don't need to generate to evaluate

· 6 min read

Every evaluation tool starts by running a model. You give it a prompt, it generates an answer, and then it scores what came out. That is the right shape when the thing you want to judge does not exist yet. It is the wrong shape when you already wrote the thing.

Most of the time, you have the text. A tagline is sitting in a doc. A support reply is drafted and waiting to send. A clause is written and you need to know if it holds up. You do not want a model to generate anything. You want a judgment on the words you already have.

Scoring is a new run type that does exactly that. It judges supplied text against a rubric or your compliance documents, and it runs no model. There is no generation step, so there is no cost, no latency, and no model in the result. It scores what you paste, returns a score per criterion, and calls a verdict. It sits alongside comparison, evaluation, and benchmarking as a first-class run type, with its own tile, its own tool, and its own result view.

“Is this compliant?”

Here is the use case that proves the point, run for real against our own marketing compliance standard. Start with a line of product copy that looks completely fine at a glance.

The copy under review

“LLM Prover gives you the most accurate model scoring available, with RAG context and compliance checks built in, so you always know your AI is right.”

Read that as a human and it reads like confident marketing. Score it against a compliance rubric backed by our regulation documents, and the picture sharpens. The rubric checks four clauses: no unbacked superlatives, feature availability stated per tier, accurate pricing, and no fabricated claims.

The verdict (per clause, with the regulation each was judged against)

  • Superlatives: fail. “most accurate” and “always” are absolute performance claims with no cited benchmark data (Clause 1.3).
  • Tier clarity: fail. RAG context and compliance scoring are gated features, and the copy names no tier (Clause 1.1).
  • Pricing: pass. No pricing figures stated.
  • Fabrication: pass. No testimonials or invented claims.

The copy scored 50 out of 100. Two flags, each tied to a specific clause, each quoting the exact phrase at fault. This is the output a marketer actually needs: not “looks fine” but “here are the two things that will not survive review, and here is the rule each one breaks.”

The fix follows directly from the flags. Drop the superlatives, state the tier.

The revised copy

“LLM Prover blends heuristic and LLM-judge scoring to rate model output against your rubric. RAG context and compliance scoring are available on Pro and above.”

Scored again, the revised line comes back at 100 out of 100, all four clauses clear. Same rubric, same regulation documents, a one-line edit between them.

Compliance score, same copy, before and after one fix (out of 100)

Per-clause scores for the same copy, before and after a single edit

Insight: Scoring judged the words you wrote, not a model’s paraphrase of them. That is the difference between generate-then-judge and scoring supplied text: when the superlative is in your copy, the judge catches it, because your copy is what it reads.

“Which phrasing is best?”

The second use case: stack-rank several versions of the same thing. Paste two taglines, three clause drafts, four support replies, and each one is scored independently against the same standard, then a winner is called. Here are four taglines we scored against a brand-voice rubric (lead with the reader’s outcome, be concrete, avoid hype, read cleanly):

The four taglines scored

  1. “LLM Prover Score: rubric-based and compliance scoring for text artifacts, with per-criterion results.” (product-led, feature list)
  2. “Judge the copy you already wrote, against your own rubric, in seconds.” (benefit-led, plain)
  3. “Score the copy you already have against your own rubric, and see exactly which lines fall short.” (benefit-led, more detail)
  4. “The most powerful way ever to check if your writing is good enough to ship.” (hype)

Four taglines, scored against one brand-voice rubric (out of 100)

Four taglines, one rubric, scored independently

The ranking is useful, but the more useful thing is what the per-criterion detail showed. The benefit-led line “Score the copy you already have against your own rubric, and see exactly which lines fall short” scored lower than expected. The breakdown explained why: the rubric’s “concrete and specific” criterion had been written in a way that rewarded naming product features and quietly penalised a line that led with the reader’s benefit instead. The rubric was pulling against itself, rewarding one thing in the “outcome” criterion and the opposite in the “concrete” criterion.

That is a problem with the standard, not the copy. The rubric was rewritten so that “concrete” means vivid and specific language, whether it describes the reader’s action or the product. Scored again, the ranking settled into something sensible, and the benefit-led line rose to match.

Insight: A stack-rank does not just rank your writing. When the ranking surprises you, it is often the rubric that needs the look, not the copy. Scoring sharpens the standard itself, so you can trust the verdicts it gives.

Score the copy you already have

Scoring is on Pro and above. Compliance scoring against your own regulation documents is on Pro Plus and above.

Get started on Pro

No model, no cost, nothing faked

The result view for a scoring run shows no cost, no latency, and no model. That is deliberate. Nothing was generated, so there is nothing to price or time.

What decides the verdict is your standard. Provide a quality rubric and the judge scores your text against your criteria. Provide a compliance store and it judges policy conformance against your regulation documents. Same text, different standard, and the standard is yours to define.

One honest boundary on the compliance use case: a compliance score is an indicator, not a certification. It tells you where copy conflicts with the clauses you supplied, which is a strong signal for a human reviewer, not a stamp of legal approval. Read the per-clause reasoning, that is where the value is.

Where to reach it

Scoring, like comparisons and evaluations, works anywhere you do:

  • Dashboard: a new scoring tile. Paste one artifact, or several to stack-rank, pick a rubric or compliance store, and run.
  • Agent and MCP: score is a core tool, so “check if this copy is compliant” or “score these three taglines against our brand rubric” works directly from your agent in Slack or your IDE.
  • API: POST /score, the same async pattern as the other run types.

Scoring is available from Pro and up, with compliance scoring on Pro Plus.

The fastest path from “I wrote this” to “is it good, is it compliant, which version wins” no longer runs through a model. You already have the text. Now you can judge it.

What’s next

evaluation compliance mcp score