Skip to content

MCP Recipes

Pro+

Check a Response Against Your Regulations

Score a response against a regulatory document using a compliance rubric and a compliance store, so you see per-clause pass/fail with the regulation text the judge used -- not a generic quality score.

Run it: paste into your MCP agent

Highlighted parts are placeholders. Replace them with your own values before (or after) copying.

Using the LLM Prover MCP tools, check whether a response complies with my regulations.
Compliance scoring needs TWO things the judge reads together: a compliance STORE (the
regulation document) and a compliance RUBRIC (criteria mapped to the clauses). Compliance
scoring is tier-gated; if my plan does not allow it the server will say so -- surface that
cleanly rather than guessing.

1. Confirm the tools you need are available (list the tools).
2. Find the compliance store: list_rag_stores(type="compliance"). It MUST have
   sync_status "ready" before it can be used -- this is a COST SAFETY gate: querying a
   store that is not ready burns a paid call to retrieve broken or empty context. If it is
   "syncing", tell me to wait and do NOT run the evaluation yet. If it is "dirty", STOP --
   do not run against a dirty store (ingestion failed/paused; the regs are incomplete and
   the verdict would be meaningless at real cost); tell me to fix the store first. Only
   proceed when sync_status is "ready". If no compliance store exists, tell me I need to
   create one and ingest the regulation document first (that is a setup step -- offer to
   walk me through it) and stop here. Do NOT fall back to a non-compliance store.
3. Find the compliance rubric: list_rubrics and pick the one with rubric_type "compliance"
   whose criteria map to the clauses I care about. If several, ask me which. If none
   exists, offer to help me build one (criteria should name the specific clauses/rules to
   check) and stop until it exists.
4. Confirm with me the prompt (and, if I am checking a specific response rather than a
   freshly generated one, the exact text) to assess, and which model(s) to run. List
   models and respect my plan's cap.
5. Run the compliance evaluation: run_evaluation(prompt=<the prompt>, models=<model(s)>,
   rubric_id=<the compliance rubric>, compliance_store_id=<the compliance store>). BOTH
   rubric_id AND compliance_store_id are required for compliance scoring -- the rubric
   defines what to check, the store supplies the regulation text the judge scores against.
   This is async -- it returns a JOB id; poll_job every 5 seconds until complete (stop
   after 5 minutes), then read the result.
6. Report per-criterion, as COMPLIANCE not quality. For each criterion the result's
   breakdown carries judge_scores (per-criterion score), judge_reasoning (why), and
   judge_criterion_chunks (the ACTUAL regulation text the judge retrieved for that
   criterion). Use judge_criterion_chunks to quote the specific clause -- do not paraphrase
   the regulation from memory. For each criterion say pass / fail / could-not-verify, give
   the reasoning, and cite the clause. Note the difference between two kinds of non-fail:
   (a) "not triggered" -- the clause did not apply to this response (e.g. no pricing stated,
   so the pricing clause is not engaged) is a PASS; (b) "no source" -- the judge found no
   matching clause in the store at all -- is NOT a pass: surface it as "could not verify"
   and flag that the store may be missing that regulation.
7. Give me a plain verdict: does the response comply overall, and exactly where it does
   not, with the offending clause named. Do not reduce compliance to a single quality
   number -- the per-clause breakdown is the point.

ON ANY FAILURE or clear mismatch (store not ready, no compliance rubric, the judge
returns nothing, every criterion comes back no-source): tell me honestly what happened and
at which step. Never report "compliant" when you could not actually score against the
regulations, and never pass a criterion the judge could not find a source clause for.

Goal

Answer “does this response comply with our regulations?” with evidence, not a vibe. By pairing a compliance rubric (what to check) with a compliance store (the regulation text), the judge scores the response clause by clause and tells you where it passes, where it fails, and which regulation it checked against. The output is a compliance verdict with citations, not a generic quality score.

When to use

Reach for this whenever “good” is defined by rules you hold as a document – marketing compliance, refund-policy adherence, a regulatory standard. It is the core compliance differentiator: the judge checks the actual response against your actual regulations, clause by clause. Pairs with drift monitoring to watch compliance over time, especially against your own deployed system after each change.

How it works

  • Two inputs, read together. A compliance evaluation needs both a compliance rubric (rubric_type: compliance, criteria mapped to clauses) and a compliance store (store_type: compliance, the regulation document). The rubric says what to check; the store supplies the text the judge scores against. Pass both to run_evaluation as rubric_id + compliance_store_id.
  • Per-clause, with citations. The result scores each criterion and surfaces the regulation clause the judge retrieved – so a failure points at the specific rule broken, and you can trust the verdict because you can see its source.
  • No source is not a pass. If the judge finds no matching clause for a criterion, that is “could not verify”, not compliance. The recipe surfaces it as a gap in the store, never as a green light.

Verification

  • The store used was store_type: compliance and sync_status: ready; a non-compliance or not-ready store was refused, not silently substituted.
  • The evaluation passed both a compliance rubric_id and compliance_store_id.
  • The report is per-criterion pass/fail with reasoning and, where available, the cited clause – not a single quality number.
  • Any “no source” criterion is reported as “could not verify”, never as a pass.

Where this leads

Natural next recipes once you are comfortable with this one.

Recipes that lead here