Compliance scoring that cites the clause, not a vibe
In a regulated setting, a quality score of 72 is not useful. What a compliance officer, an auditor, or a risk owner needs is different: which specific rule did this output break, and what does the rule say. A single blended number cannot answer that. Clause-by-clause scoring can.
This is a real run. An AI model drafted a piece of marketing copy, and before it could ship, a compliance evaluation scored it against an actual regulatory document, returning a pass or fail for each clause with the regulation text attached.
The setup
Three inputs define the run:
- the brief the copywriting model was given
- the regulatory document, held in a compliance store
- the compliance rubric that maps each criterion to a clause
The brief given to the model
The brief deliberately asks for things a marketing-standards document tends to prohibit: an absolute superlative and a customer quote. The compliance rubric scores against four clauses of the real regulatory document held in a compliance store.
The compliance rubric (criterion to clause)
- Availability: feature availability stated per tier (Clause 1.1)
- Superlatives: no absolute performance superlatives (Clause 1.3)
- Pricing: pricing figures accurate (Clause 1.2 / 3.2)
- Fabrication: no fabricated testimonials presented as real (Clause 6.1)
What the model wrote
It reads like normal marketing copy. That is the point: nothing here looks obviously wrong at a glance, which is exactly why a human reviewer under time pressure waves it through.
The verdict, clause by clause
The compliance evaluation scored it 50 out of 100 and broke the result down by clause.
| Criterion | Clause | Verdict |
|---|---|---|
| Superlatives | 1.3 | fail |
| Fabrication | 6.1 | fail |
| Pricing | 1.2 / 3.2 | pass |
| Availability | 1.1 | pass |
Two failures, each tied to a specific clause and a specific phrase. The superlative clause failed on “the fastest and most accurate benchmarking platform”, an absolute performance claim with no cited benchmark data. The fabrication clause failed on the customer quote, an unattributed testimonial presented as real.
The part that makes this auditable rather than a judgment call: the evaluation returns the actual regulation text the judge applied to each clause. For the superlative failure, it carried the rule verbatim:
A reviewer does not have to trust the score. They can read the exact rule the output broke, see the phrase that broke it, and act. That is the difference between a compliance result and a confidence number.
Score your outputs against your own regulations
Compliance scoring and compliance stores are on Pro Plus. Clause-by-clause, with the regulation text attached.
The failure was the brief, not the model
The evaluation did not stop at pass or fail. The diagnostic traced both violations to their source, and in this run the source was the prompt: it had been instructed to write a superlative and to include a quote. The model did what it was told. The diagnostic said so, pointing at the brief rather than blaming the model.
That matters for a content pipeline. When a compliance check fails, the next question is always “where do we fix it”, and the honest answer is often upstream of the model: the brief, the template, the house style that tells writers to reach for superlatives. The diagnostic names the layer so the fix lands in the right place. The same layered diagnosis runs across every stage, as shown in how the diagnostic improvement loop works.
Why this is the regulated-industry case
For legal, financial, healthcare, and insurance teams, AI output that must align with regulations is a standing risk. A model update, a new template, or a drifted prompt can quietly start producing non-compliant output, and the first sign is usually an auditor or a regulator, not a dashboard. A scheduled compliance benchmark catches the change first, clause by clause, with the regulation text on record.
Compliance scoring and compliance stores are Pro Plus features. The full framework of levers this fits into, including how compliance interacts with model choice, store content, and drift, is in the practical framework for your LLM stack.
What’s next
Compliance Store Guide
Upload regulations or policies as a compliance store and score model outputs against specific clauses.
Scoring Guide
How LLM Prover calculates scores, what affects them, and how to use the diagnostic judge to find out why.
RAG Store Guide
How to create a context store, upload documents, and get reliable retrieval in comparisons, evaluations, and benchmarks.