Skip to content

MCP Recipes

Pro

Score a Model Against a Known Answer

Give a prompt and the ideal answer; the agent runs an evaluation and tells you how close a model got, with a quality score and the judge's reasoning. The gentlest way to start scoring.

Run it: paste into your MCP agent

Highlighted parts are placeholders. Replace them with your own values before (or after) copying.

Using the LLM Prover MCP tools, score how well a model answers a task with a known-good
answer. This is the simplest form of scoring -- a gold-standard comparison, no rubric.

1. Confirm the tools you need are available (list the tools).
2. Ask me for the prompt and the ideal (gold-standard) answer to score against, if I have
   not already given them.
3. Pick one model to start with -- list_models and choose a sensible default, or ask me if
   I have a preference. Keep it to one model for a first evaluation.
4. Run the evaluation: run_evaluation(prompt=<the prompt>, models=<the one model>,
   expected_output=<the gold-standard answer>). This is async -- it returns a JOB id;
   poll_job every 5 seconds until complete (stop after 5 minutes), then read the result.
5. Report in plain language: the quality score, and WHY the judge scored it that way (the
   judge's reasoning and the per-dimension breakdown). If the score is lower than I might
   expect, explain what the response missed versus the gold-standard answer.
6. Offer the natural next step: scoring on richer criteria (a rubric) or comparing several
   models to find the best value -- but only mention it once, and only if it fits.

ON ANY FAILURE or clear mismatch (the eval will not score, the model errored, the judge
returns nothing): tell me honestly what happened and at which step -- do not invent a score.

Goal

Take the first step into scoring with zero setup: a prompt, the ideal answer, one model, and a number that tells you how close the model got – plus the judge’s reasoning for why. No rubric to write, no benchmark to configure.

When to use

Reach for this the first time you want to measure quality rather than just read output, or any time you have a task with a clear known-good answer and want a quick read on a model. It is the gentlest on-ramp to everything else: rubrics, model procurement, drift monitoring.

How it works

  • Gold standard in, score out. You supply the ideal answer; the evaluation scores the model’s response against it and returns a quality score with the judge’s reasoning and a per-dimension breakdown (coverage, precision, conciseness, structure).
  • One model, one prompt. Deliberately minimal – this is the starting point. Comparing several models or scoring on weighted criteria are the next steps, not this one.

Verification

  • An evaluation ran against the gold-standard answer and returned a quality score.
  • The report includes the judge’s reasoning, not just the number.
  • On failure, that is stated plainly – no invented score.

Where this leads

Natural next recipes once you are comfortable with this one.