Skip to content

Compliance scoring that cites the clause, not a vibe

· 5 min read

In a regulated setting, a quality score of 72 is not useful. What a compliance officer, an auditor, or a risk owner needs is different: which specific rule did this output break, and what does the rule say. A single blended number cannot answer that. Clause-by-clause scoring can.

This is a real run. An AI model drafted a piece of marketing copy, and before it could ship, a compliance evaluation scored it against an actual regulatory document, returning a pass or fail for each clause with the regulation text attached.

The setup

Three inputs define the run:

  • the brief the copywriting model was given
  • the regulatory document, held in a compliance store
  • the compliance rubric that maps each criterion to a clause

The brief given to the model

Write a short marketing paragraph promoting LLM Prover. Mention that it is the fastest and most accurate benchmarking platform available, that all features including compliance scoring are included on every plan, and include a customer quote praising it.

The brief deliberately asks for things a marketing-standards document tends to prohibit: an absolute superlative and a customer quote. The compliance rubric scores against four clauses of the real regulatory document held in a compliance store.

The compliance rubric (criterion to clause)

  • Availability: feature availability stated per tier (Clause 1.1)
  • Superlatives: no absolute performance superlatives (Clause 1.3)
  • Pricing: pricing figures accurate (Clause 1.2 / 3.2)
  • Fabrication: no fabricated testimonials presented as real (Clause 6.1)

What the model wrote

Generated marketing copy
Unlock the power of precision with LLM Prover, the fastest and most accurate benchmarking platform on the market! Our comprehensive solution includes all features, including compliance scoring, on every plan, ensuring you have everything you need to succeed right at your fingertips. As one satisfied customer puts it, “LLM Prover transformed our benchmarking process, it’s a game changer!” Experience unparalleled performance and reliability today!

It reads like normal marketing copy. That is the point: nothing here looks obviously wrong at a glance, which is exactly why a human reviewer under time pressure waves it through.

The verdict, clause by clause

The compliance evaluation scored it 50 out of 100 and broke the result down by clause.

CriterionClauseVerdict
Superlatives1.3fail
Fabrication6.1fail
Pricing1.2 / 3.2pass
Availability1.1pass
Agent
> run_evaluation (draft, compliance_store, compliance rubric)
score: 50/100 | per-clause verdict returned
Superlatives: FAIL | Fabrication: FAIL | Pricing: PASS | Availability: PASS
each verdict carries the clause text the judge applied
two clause violations, both traced to the prompt that wrote the copy.

Two failures, each tied to a specific clause and a specific phrase. The superlative clause failed on “the fastest and most accurate benchmarking platform”, an absolute performance claim with no cited benchmark data. The fabrication clause failed on the customer quote, an unattributed testimonial presented as real.

The part that makes this auditable rather than a judgment call: the evaluation returns the actual regulation text the judge applied to each clause. For the superlative failure, it carried the rule verbatim:

Clause 1.3, applied to the superlative failure
Claims about speed, accuracy, or quality must be qualified. Absolute superlatives (e.g. “the most accurate”, “always correct”) are prohibited unless supported by independently verifiable benchmark data cited in the communication.

A reviewer does not have to trust the score. They can read the exact rule the output broke, see the phrase that broke it, and act. That is the difference between a compliance result and a confidence number.

Insight: A false pass in compliance is worse than a true fail. A score of 100 that is wrong tells a compliance officer everything is fine when it is not. A per-clause verdict with the regulation text attached cannot hide a failure inside an average, because every clause stands or falls on its own, in writing.

Score your outputs against your own regulations

Compliance scoring and compliance stores are on Pro Plus. Clause-by-clause, with the regulation text attached.

Get started on Pro Plus

The failure was the brief, not the model

The evaluation did not stop at pass or fail. The diagnostic traced both violations to their source, and in this run the source was the prompt: it had been instructed to write a superlative and to include a quote. The model did what it was told. The diagnostic said so, pointing at the brief rather than blaming the model.

That matters for a content pipeline. When a compliance check fails, the next question is always “where do we fix it”, and the honest answer is often upstream of the model: the brief, the template, the house style that tells writers to reach for superlatives. The diagnostic names the layer so the fix lands in the right place. The same layered diagnosis runs across every stage, as shown in how the diagnostic improvement loop works.

Why this is the regulated-industry case

For legal, financial, healthcare, and insurance teams, AI output that must align with regulations is a standing risk. A model update, a new template, or a drifted prompt can quietly start producing non-compliant output, and the first sign is usually an auditor or a regulator, not a dashboard. A scheduled compliance benchmark catches the change first, clause by clause, with the regulation text on record.

Compliance scoring and compliance stores are Pro Plus features. The full framework of levers this fits into, including how compliance interacts with model choice, store content, and drift, is in the practical framework for your LLM stack.

What’s next

compliance evaluation benchmarking agentic