MCP Recipes
ProBuild a Scoring Rubric From a Description
Describe what a good answer looks like in plain language; the agent drafts a weighted rubric, trials it on a real response, and saves a reusable rubric you can score evaluations and benchmarks against.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, help me turn a plain-language description of a good answer into a reusable scoring rubric, and tune it on a real response before saving. 1. Confirm the tools you need are available (list the tools). 2. Ask me two things if I have not already said them: (a) a plain-language description of what makes a good answer for my task, and (b) a sample prompt to trial the rubric on. A sample ideal answer is helpful but optional. 3. Draft the rubric criteria from my description. Each criterion needs a short name, a relative weight (a positive number -- weights are relative, they do NOT need to sum to 100), and a SPECIFIC description of what the judge should check (vague descriptions produce inconsistent scores). Respect my plan's criteria-per-rubric cap -- if you are unsure, keep it to a handful and tell me I can add more on a higher tier. Show me the drafted criteria and let me adjust before you save anything. 4. Create the rubric: create_rubric(name, criteria[, description]). Note the rubric_id. 5. Trial it on my sample: run_evaluation(prompt=<my sample prompt>, models=<one cheap model you pick from list_models>, rubric_id=<the new rubric_id>). This is async -- it returns a JOB id; poll_job every 5 seconds until complete (stop after 5 minutes), then read the result from the job (or via list_evaluations). 6. Show me the per-criterion scores AND the judge's reasoning for each criterion. Also read the result's pipeline_diagnostics findings: they flag when a criterion is miscalibrated against the prompt (e.g. "criterion penalises format the prompt never asked to avoid"). These findings are often the clearest signal that the rubric -- not the model -- needs adjusting. This is the point of the trial: we are checking the rubric behaves, not grading the model. 7. Refine if needed: if a criterion scored in a way that does not match my intent (too harsh, too lenient, judging something I did not ask for, or flagged as mismatched in the diagnostics), propose a wording or weight change and, on my OK, call update_rubric. TWO things to get right about update_rubric: (a) it requires name AND the FULL criteria list every time -- it replaces the whole criteria set, so always send the complete updated list, not just the one you changed; (b) it only persists the fields you send, so ALSO resend description and rubric_type, or they get nulled. Re-trial (step 5) until the rubric scores the way I expect. 8. Hand back the final rubric_id and remind me I can now attach it to any evaluation or benchmark via rubric_id -- it is reusable, not single-use. ON ANY FAILURE or clear mismatch (a tool errors, the eval will not score, the judge returns nothing): tell me honestly what happened and at which step -- do not invent scores or claim the rubric is good when you could not actually trial it. If the rubric was created but the trial failed, tell me the rubric_id still exists so we do not create duplicates on retry.
Goal
Remove the hardest part of scoring: writing good rubric criteria. You say what “good” means in plain language; the agent turns that into weighted, specifically-worded criteria, proves they behave by scoring a real response, tunes them with you, and leaves you with a saved rubric you can reuse across every evaluation and benchmark.
When to use
Reach for this the first time you want rubric-based scoring and don’t want to hand-craft criteria, or whenever your current rubric isn’t scoring the way you expect and you want to tune it against real output. A good rubric is reusable infrastructure – build it once, score with it everywhere.
How it works
- Description in, criteria out. The agent translates your plain-language “good answer” into named, weighted criteria with specific judge instructions. Weights are relative, so you express priority (“accuracy matters twice as much as tone”) without doing arithmetic.
- Trial before trust. The rubric is scored against a real response via an evaluation, so you see the per-criterion scores and the judge’s reasoning before you rely on it. The trial grades the rubric, not the model – and the evaluation’s pipeline diagnostics will explicitly flag a criterion that is miscalibrated against the prompt, which is the fastest way to spot a rubric that needs tuning.
- Tune, then reuse.
update_rubricreplaces the full criteria set, so refinement is a loop: adjust wording or weights, re-trial, repeat until it behaves. The result is a savedrubric_idyou attach to any evaluation or benchmark.
Verification
- A rubric exists (
rubric_idreturned) with the criteria you approved. - A trial evaluation ran against it and surfaced per-criterion scores + reasoning.
- Any refinement replaced the full criteria list (not a partial update) and was re-trialled.
- The final
rubric_idis reported back and is reusable inrun_evaluation/create_benchmarkviarubric_id.
Where this leads
Natural next recipes once you are comfortable with this one.
Recipes that lead here