MCP Recipes
Pro+Check a Response Against Your Regulations
Score a response against a regulatory document using a compliance rubric and a compliance store, so you see per-clause pass/fail with the regulation text the judge used -- not a generic quality score.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, check whether a response complies with my regulations. Compliance scoring needs TWO things the judge reads together: a compliance STORE (the regulation document) and a compliance RUBRIC (criteria mapped to the clauses). Compliance scoring is tier-gated; if my plan does not allow it the server will say so -- surface that cleanly rather than guessing. 1. Confirm the tools you need are available (list the tools). 2. Find the compliance store: list_rag_stores(type="compliance"). It MUST have sync_status "ready" before it can be used -- this is a COST SAFETY gate: querying a store that is not ready burns a paid call to retrieve broken or empty context. If it is "syncing", tell me to wait and do NOT run the evaluation yet. If it is "dirty", STOP -- do not run against a dirty store (ingestion failed/paused; the regs are incomplete and the verdict would be meaningless at real cost); tell me to fix the store first. Only proceed when sync_status is "ready". If no compliance store exists, tell me I need to create one and ingest the regulation document first (that is a setup step -- offer to walk me through it) and stop here. Do NOT fall back to a non-compliance store. 3. Find the compliance rubric: list_rubrics and pick the one with rubric_type "compliance" whose criteria map to the clauses I care about. If several, ask me which. If none exists, offer to help me build one (criteria should name the specific clauses/rules to check) and stop until it exists. 4. Confirm with me the prompt (and, if I am checking a specific response rather than a freshly generated one, the exact text) to assess, and which model(s) to run. List models and respect my plan's cap. 5. Run the compliance evaluation: run_evaluation(prompt=<the prompt>, models=<model(s)>, rubric_id=<the compliance rubric>, compliance_store_id=<the compliance store>). BOTH rubric_id AND compliance_store_id are required for compliance scoring -- the rubric defines what to check, the store supplies the regulation text the judge scores against. This is async -- it returns a JOB id; poll_job every 5 seconds until complete (stop after 5 minutes), then read the result. 6. Report per-criterion, as COMPLIANCE not quality. For each criterion the result's breakdown carries judge_scores (per-criterion score), judge_reasoning (why), and judge_criterion_chunks (the ACTUAL regulation text the judge retrieved for that criterion). Use judge_criterion_chunks to quote the specific clause -- do not paraphrase the regulation from memory. For each criterion say pass / fail / could-not-verify, give the reasoning, and cite the clause. Note the difference between two kinds of non-fail: (a) "not triggered" -- the clause did not apply to this response (e.g. no pricing stated, so the pricing clause is not engaged) is a PASS; (b) "no source" -- the judge found no matching clause in the store at all -- is NOT a pass: surface it as "could not verify" and flag that the store may be missing that regulation. 7. Give me a plain verdict: does the response comply overall, and exactly where it does not, with the offending clause named. Do not reduce compliance to a single quality number -- the per-clause breakdown is the point. ON ANY FAILURE or clear mismatch (store not ready, no compliance rubric, the judge returns nothing, every criterion comes back no-source): tell me honestly what happened and at which step. Never report "compliant" when you could not actually score against the regulations, and never pass a criterion the judge could not find a source clause for.
Goal
Answer “does this response comply with our regulations?” with evidence, not a vibe. By pairing a compliance rubric (what to check) with a compliance store (the regulation text), the judge scores the response clause by clause and tells you where it passes, where it fails, and which regulation it checked against. The output is a compliance verdict with citations, not a generic quality score.
When to use
Reach for this whenever “good” is defined by rules you hold as a document – marketing compliance, refund-policy adherence, a regulatory standard. It is the core compliance differentiator: the judge checks the actual response against your actual regulations, clause by clause. Pairs with drift monitoring to watch compliance over time, especially against your own deployed system after each change.
How it works
- Two inputs, read together. A compliance evaluation needs both a compliance rubric
(
rubric_type: compliance, criteria mapped to clauses) and a compliance store (store_type: compliance, the regulation document). The rubric says what to check; the store supplies the text the judge scores against. Pass both torun_evaluationasrubric_id+compliance_store_id. - Per-clause, with citations. The result scores each criterion and surfaces the regulation clause the judge retrieved – so a failure points at the specific rule broken, and you can trust the verdict because you can see its source.
- No source is not a pass. If the judge finds no matching clause for a criterion, that is “could not verify”, not compliance. The recipe surfaces it as a gap in the store, never as a green light.
Verification
- The store used was
store_type: complianceandsync_status: ready; a non-compliance or not-ready store was refused, not silently substituted. - The evaluation passed both a compliance
rubric_idandcompliance_store_id. - The report is per-criterion pass/fail with reasoning and, where available, the cited clause – not a single quality number.
- Any “no source” criterion is reported as “could not verify”, never as a pass.
Where this leads
Natural next recipes once you are comfortable with this one.
Recipes that lead here