Skip to content

MCP Recipes

Pro

Build a Scoring Rubric From a Description

Describe what a good answer looks like in plain language; the agent drafts a weighted rubric, trials it on a real response, and saves a reusable rubric you can score evaluations and benchmarks against.

Run it: paste into your MCP agent

Highlighted parts are placeholders. Replace them with your own values before (or after) copying.

Using the LLM Prover MCP tools, help me turn a plain-language description of a good
answer into a reusable scoring rubric, and tune it on a real response before saving.

1. Confirm the tools you need are available (list the tools).
2. Ask me two things if I have not already said them: (a) a plain-language description
   of what makes a good answer for my task, and (b) a sample prompt to trial the rubric
   on. A sample ideal answer is helpful but optional.
3. Draft the rubric criteria from my description. Each criterion needs a short name, a
   relative weight (a positive number -- weights are relative, they do NOT need to sum to
   100), and a SPECIFIC description of what the judge should check (vague descriptions
   produce inconsistent scores). Respect my plan's criteria-per-rubric cap -- if you are
   unsure, keep it to a handful and tell me I can add more on a higher tier. Show me the
   drafted criteria and let me adjust before you save anything.
4. Create the rubric: create_rubric(name, criteria[, description]). Note the rubric_id.
5. Trial it on my sample: run_evaluation(prompt=<my sample prompt>, models=<one cheap
   model you pick from list_models>, rubric_id=<the new rubric_id>). This is async -- it
   returns a JOB id; poll_job every 5 seconds until complete (stop after 5 minutes), then
   read the result from the job (or via list_evaluations).
6. Show me the per-criterion scores AND the judge's reasoning for each criterion. Also
   read the result's pipeline_diagnostics findings: they flag when a criterion is
   miscalibrated against the prompt (e.g. "criterion penalises format the prompt never
   asked to avoid"). These findings are often the clearest signal that the rubric -- not
   the model -- needs adjusting. This is the point of the trial: we are checking the
   rubric behaves, not grading the model.
7. Refine if needed: if a criterion scored in a way that does not match my intent (too
   harsh, too lenient, judging something I did not ask for, or flagged as mismatched in
   the diagnostics), propose a wording or weight change and, on my OK, call update_rubric.
   TWO things to get right about update_rubric: (a) it requires name AND the FULL criteria
   list every time -- it replaces the whole criteria set, so always send the complete
   updated list, not just the one you changed; (b) it only persists the fields you send,
   so ALSO resend description and rubric_type, or they get nulled. Re-trial (step 5) until
   the rubric scores the way I expect.
8. Hand back the final rubric_id and remind me I can now attach it to any evaluation or
   benchmark via rubric_id -- it is reusable, not single-use.

ON ANY FAILURE or clear mismatch (a tool errors, the eval will not score, the judge
returns nothing): tell me honestly what happened and at which step -- do not invent scores
or claim the rubric is good when you could not actually trial it. If the rubric was created
but the trial failed, tell me the rubric_id still exists so we do not create duplicates on
retry.

Goal

Remove the hardest part of scoring: writing good rubric criteria. You say what “good” means in plain language; the agent turns that into weighted, specifically-worded criteria, proves they behave by scoring a real response, tunes them with you, and leaves you with a saved rubric you can reuse across every evaluation and benchmark.

When to use

Reach for this the first time you want rubric-based scoring and don’t want to hand-craft criteria, or whenever your current rubric isn’t scoring the way you expect and you want to tune it against real output. A good rubric is reusable infrastructure – build it once, score with it everywhere.

How it works

  • Description in, criteria out. The agent translates your plain-language “good answer” into named, weighted criteria with specific judge instructions. Weights are relative, so you express priority (“accuracy matters twice as much as tone”) without doing arithmetic.
  • Trial before trust. The rubric is scored against a real response via an evaluation, so you see the per-criterion scores and the judge’s reasoning before you rely on it. The trial grades the rubric, not the model – and the evaluation’s pipeline diagnostics will explicitly flag a criterion that is miscalibrated against the prompt, which is the fastest way to spot a rubric that needs tuning.
  • Tune, then reuse. update_rubric replaces the full criteria set, so refinement is a loop: adjust wording or weights, re-trial, repeat until it behaves. The result is a saved rubric_id you attach to any evaluation or benchmark.

Verification

  • A rubric exists (rubric_id returned) with the criteria you approved.
  • A trial evaluation ran against it and surfaced per-criterion scores + reasoning.
  • Any refinement replaced the full criteria list (not a partial update) and was re-trialled.
  • The final rubric_id is reported back and is reusable in run_evaluation / create_benchmark via rubric_id.

Where this leads

Natural next recipes once you are comfortable with this one.

Recipes that lead here