MCP Recipes
Pro+Score Existing Copy for Compliance
Judge a piece of text you already have (marketing copy, a policy line, a drafted answer) against a compliance rubric and your regulation store, with no model run, so you get a per-clause verdict on text you hand it, before it ships.
Agentic and programmatic use. Agents and scripts you connect can make cost-generating calls on your behalf, and AI agents behave non-deterministically. You are responsible for monitoring and bounding your own automation. See Terms for details.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, score a piece of text I ALREADY HAVE against my
regulations. This is the Score run type: no model is run, the judge scores exactly the
text I supply. Compliance scoring is tier-gated (Pro+); if my plan does not allow it the
server will say so, so surface that cleanly rather than guessing.
1. Confirm the tools you need are available (list the tools). The tool is `score`. It is a
core tool, so it should be in your list; if it is not, find it via list_tool_catalog,
get its schema with request_tools, then call it through invoke("score", {...}).
2. Find the compliance store: list_rag_stores(type="compliance"). It MUST have
sync_status "ready" before it can be used. This is a cost-safety gate: scoring against
a store that is not ready would judge against broken or empty regulation text. If it is
"syncing", tell me to wait and do NOT run yet. If it is "dirty", STOP and tell me to fix
the store first. If no compliance store exists, tell me I need to create one and ingest
the regulation document first (offer to walk me through it) and stop. Do NOT fall back
to a non-compliance store.
3. Find the compliance rubric: list_rubrics and pick the one with rubric_type "compliance"
whose criteria map to the clauses I care about. If several, ask me which. If none
exists, offer to help me build one and stop until it exists.
4. Confirm with me the exact text to score (the artifact). Do not paraphrase or clean it
up; score what I actually give you.
5. Run the score: score(artifact=<the exact text>, rubric_id=<the compliance rubric>,
compliance_store_id=<the compliance store>). Do NOT pass models, providers, params, or
any benchmark setting; Score runs no model and rejects them. This is async: it returns a
JOB id; poll_job every 5 seconds until complete (stop after 5 minutes), then read the
result.
6. Report per-criterion, as COMPLIANCE not quality. For each criterion the result's
breakdown carries judge_scores (per-criterion score), judge_reasoning (why), and
judge_criterion_chunks (the ACTUAL regulation text the judge used). Quote the clause
from judge_criterion_chunks; do not paraphrase the regulation from memory. For each
criterion say pass / fail / could-not-verify, give the reasoning, and cite the clause.
Distinguish "not triggered" (the clause did not apply, e.g. no pricing stated so the
pricing clause is not engaged, which is a PASS) from "no source" (the judge found no
matching clause in the store at all, which is NOT a pass: report it as "could not
verify" and flag that the store may be missing that regulation).
7. Give me a plain verdict: does my text comply overall, and exactly where it does not,
with the offending clause named. The per-clause breakdown is the point; do not reduce it
to one number. There is no cost or latency to report, because no model ran.
ON ANY FAILURE or clear mismatch (store not ready, no compliance rubric, the judge returns
nothing, every criterion comes back no-source): tell me honestly what happened and at which
step. Never report "compliant" when you could not actually score against the regulations.
Goal
Answer “does this copy comply with our rules?” for text you already wrote, before it ships. Score judges a supplied artifact (a tagline, a policy line, a support reply) against a compliance rubric and your regulation store, clause by clause, and tells you where it passes, where it fails, and which regulation it checked against. No model generates anything; the judge scores exactly what you hand it.
When to use
Reach for this when you have the text in hand and need a compliance verdict before it goes out: marketing copy against your AUP, a refund-policy statement against the policy document, a drafted clause against a standard. If instead you want to check a response your MODEL produced, use the compliance evaluation recipe, which generates then scores.
How it works
- Supplied text, no generation. Score takes your artifact as the thing under test, where a model response would otherwise sit. There is no model, so no cost, no latency, no model field in the result. Those are absent, not zero, because nothing was generated.
- Two inputs, read together. Compliance scoring needs both a compliance rubric
(
rubric_type: compliance, criteria mapped to clauses) and a compliance store (store_type: compliance, the regulation text). The rubric says what to check; the store supplies the text the judge scores against. Pass both toscoreasrubric_id+compliance_store_id. - Per-clause, with citations. Each criterion is scored and surfaces the regulation clause the judge used, so a failure names the specific rule broken and you can see its source.
- No source is not a pass. If the judge finds no matching clause for a criterion, that is “could not verify”, a gap in the store, never a green light.
Verification
- The run used the
scoretool with a suppliedartifact, not a generate-then-score path. - The store was
store_type: complianceandsync_status: ready; a non-compliance or not-ready store was refused, not silently substituted. - Both a compliance
rubric_idandcompliance_store_idwere passed; no model/params were. - The report is per-criterion pass/fail with reasoning and the cited clause, not a single quality number, and the result carries no cost/latency/model.
- Any “no source” criterion is reported as “could not verify”, never as a pass.
Where this leads
Natural next recipes once you are comfortable with this one.
Have a few versions of the copy? Score them all at once and let the ranking pick the strongest against the same rules.
Checking a response your MODEL produced rather than text you wrote? Run a compliance evaluation that generates then scores, instead of scoring supplied text.
Recipes that lead here