MCP Recipes
ProPick the Cheapest Model Above a Quality Bar
Score a set of candidate models on your task, then get a recommendation for the cheapest one that clears the quality bar you set -- a procurement decision, not a table to read.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, help me choose the cheapest model that is still good enough for my task. The output is a RECOMMENDATION I approve, not just a results table. 1. Confirm the tools you need are available (list the tools). 2. Discover the models available to my account (list_models) and show me the candidates. Respect my plan's model-per-comparison cap -- if I want more candidates than the cap, tell me and let me trim. Do not hardcode model names; use what list_models returns. 3. Get what we need to score quality -- a comparison alone is not enough, because "good enough" needs a number. Ask me for EITHER a gold-standard answer (expected_output) OR a saved rubric (rubric_id) for the task. If I have neither, offer to help me build a rubric first (that is a separate recipe) or to supply a gold-standard answer. 4. Confirm the quality bar with me (e.g. "at least 85 out of 100"). If I do not give one, propose a sensible default and say so explicitly -- do not silently assume. 5. Run a SCORED evaluation across all candidates in one call: run_evaluation(prompt=<task>, models=<the candidate map>, expected_output=<gold> OR rubric_id=<rubric>). This is async -- it returns a JOB id; poll_job every 5 seconds until complete (stop after 5 minutes), then read the per-model results from the job. 6. Build the decision from the scored results: for each model you have quality_score and cost_usd. Discard any model that errored or did not score. Among the models at or above my bar, pick the CHEAPEST -- that is the recommendation. NOTE: compute this yourself from quality_score + cost_usd; do NOT just echo the result's built-in winner. The "by_efficiency" winner is best quality-per-dollar, which is a DIFFERENT decision from "cheapest above the bar" -- a pricier, much-better model can win efficiency while a cheaper model clears the bar. They sometimes coincide; do not assume they do. 7. Report the decision in plain language: name the recommended model, its quality_score and cost, and WHY (cheapest above the bar). Then give me the tradeoff context: the highest-quality model regardless of cost, and the next-cheapest alternative, so I can overrule the bar if I want. If NO model clears the bar, say so plainly and show the top scorer -- do not recommend a model that failed the bar as if it passed. ON ANY FAILURE or clear mismatch (the eval will not score, every candidate errored, the judge returns nothing): tell me honestly what happened and at which step -- do not invent scores or recommend a model you could not actually score. If only some candidates failed, make the recommendation from the ones that succeeded and name which were excluded and why.
Goal
Turn “which model should I use?” into a decision with a cost constraint, made for you. You give the task, a way to score it, and a quality bar; the agent scores the candidates and recommends the cheapest model that is still good enough – with the tradeoffs spelled out so you can overrule it. The deliverable is a recommendation, not a results grid.
When to use
Reach for this at build time, when you are choosing a model for a real task and cost matters: “find me the cheapest model that scores at least 85 on this prompt.” It is the procurement decision that sits one step before you commit a model to production – and it pairs naturally with drift monitoring once you have chosen.
How it works
- Scored, not just compared. “Good enough” needs a number, so the decision runs through a scored evaluation (gold standard or rubric), not a bare comparison. Every candidate gets a quality_score the bar can be applied to.
- The decision is cheapest-above-the-bar. The agent filters to models meeting your bar, then picks the lowest cost among them. Models that errored or did not score are excluded, never silently recommended.
- Tradeoffs stay visible. The recommendation names the runner-up and the highest-quality-regardless-of-cost option, so you can consciously spend more for more quality – or relax the bar – rather than taking the pick blind.
Verification
- All candidate models came from list_models (tier-scoped) and respected the per-comparison cap.
- The decision used scored results (a quality_score per model from a gold standard or rubric).
- The recommended model is the cheapest at or above the stated bar; excluded/failed models are named, not silently dropped.
- If no model cleared the bar, that is stated plainly with the top scorer shown – no model that failed the bar is presented as a pass.
Where this leads
Natural next recipes once you are comfortable with this one.
Lock in your choice -- monitor the model you picked over time and get told if it ever drops below the bar.
Judge procurement on exactly what you care about -- build a reusable rubric and re-run the decision against it.
Recipes that lead here