MCP Recipes
ProRank Several Drafts Head to Head
Hand the agent several versions of the same thing (two taglines, three clause drafts, four support replies) and score them all against one standard, so you get each one judged and a clear winner, without running a model.
Agentic and programmatic use. Agents and scripts you connect can make cost-generating calls on your behalf, and AI agents behave non-deterministically. You are responsible for monitoring and bounding your own automation. See Terms for details.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, rank several versions of a piece of text against ONE
standard and tell me which is strongest. This is the Score run type with multiple
artifacts: no model is run, each version I supply is judged against the same standard, and
a winner is returned.
1. Confirm the tools you need are available (list the tools). The tool is `score`. It is a
core tool; if it is not in your list, find it via list_tool_catalog, get its schema with
request_tools, then call it through invoke("score", {...}).
2. Collect the versions from me: two or more distinct pieces of text to rank. Keep them as
I give them; do not edit or merge them. If I only give one, this is the single-artifact
case, so point me to the single-score flow instead.
3. Agree the ONE standard to judge them all against. Either:
- a quality rubric: list_rubrics and pick a rubric_type "quality" one whose criteria
capture what "best" means here (clarity, tone, completeness, and so on); or
- compliance: a rubric_type "compliance" rubric PLUS a compliance store
(list_rag_stores(type="compliance"), which must be sync_status "ready"), when "best"
means "most compliant with our regulations". For compliance, both are required.
If no suitable rubric exists, offer to help me build one and stop until it exists. Use
ONE standard for all artifacts; that is what makes the ranking fair.
4. Run the ranking: score(artifacts=[<version 1>, <version 2>, ...], rubric_id=<rubric>)
and, for compliance, also compliance_store_id=<store>. Do NOT pass models, providers,
params, or benchmark settings; Score runs no model and rejects them. This is async: it
returns a JOB id; poll_job every 5 seconds until complete (stop after 5 minutes).
5. Report the ranking. The result has one entry per artifact (Artifact 1, Artifact 2, ...)
each with its own scores and per-criterion breakdown, plus a winner block. Give me the
artifacts ordered best to worst with their scores, name the winner, and say briefly WHY
it won against the standard. If the result reports a tie, say so honestly rather than
breaking it arbitrarily. For a compliance standard, cite the clause behind a decisive
difference where one is available.
6. There is no cost, latency, or model to report; nothing was generated. Do not invent a
cost or a "fastest" dimension.
ON ANY FAILURE or clear mismatch (only one artifact supplied, no suitable rubric, a
compliance store that is not ready, the judge returns nothing): tell me honestly what
happened and at which step. Never declare a winner you could not actually score.
Goal
Pick the strongest of several versions you already have, judged against one standard. Hand Score two or more artifacts and a single rubric (or a compliance rubric plus store), and it scores each one the same way and names a winner. “Which phrasing is best for our ToS”, “which of these three taglines”, “which support reply best follows the SOP”: a ranking with the reasoning, not a guess.
When to use
Reach for this when you have the candidates in hand and need to choose: competing drafts, A/B copy, alternative clause wordings. The standard you pick decides what “best” means, a quality rubric judges the writing, a compliance rubric plus store judges conformance to your rules. For a yes/no verdict on a single piece of text, use the single-artifact score instead.
How it works
- Supplied versions, no generation. Each artifact you pass is scored where a model response would otherwise sit. No model runs, so the result carries no cost, latency, or model, and no “fastest” axis; ranking is on the score alone.
- One standard, applied to all. Every artifact is judged against the same rubric (and compliance store, for compliance), which is what makes the comparison fair. Mixing standards would make the ranking meaningless.
- A winner, honestly. The result returns the highest scorer as the winner and surfaces ties rather than hiding them. Each artifact keeps its own per-criterion detail, so you can see not just who won but where each one was strong or weak.
Verification
- The run used the
scoretool with anartifactslist of two or more, not a single artifact and not a generate-then-score path. - One standard was applied to every artifact (a quality rubric, or a compliance rubric plus a ready compliance store); no model/params were passed.
- The report ranks the artifacts with per-criterion detail and names the winner, with ties surfaced rather than broken arbitrarily.
- The result carries no cost/latency/model, and no invented performance dimension.
Where this leads
Natural next recipes once you are comfortable with this one.
Just need a yes/no on one piece of copy against your regulations? Score a single artifact for a per-clause compliance verdict.
Judging against the same standard a lot? Save it as a preset and run the ranking by name next time.
Recipes that lead here