MCP Recipes
pro+ onlyMonitor a Model for Drift Over Time
Stand up a scheduled benchmark, then have your agent reconcile each run against the baseline and tell you only when quality, cost, or latency moves.
Run it -- paste into your MCP agent
Highlighted parts are placeholders -- replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, set up (or reconcile) ongoing drift monitoring. This recipe keeps ONE standing note as the monitoring anchor -- unlike a one-shot run, the note stays "running" between cycles and is only closed when I end the series. FIRST, work out whether we are starting a monitor or reconciling an existing one: 1. Confirm the tools you need are available (list the tools). 2. Read your standing notes for this recipe: get_agent_notes(recipe_id="recipe_id"). - If an active note already exists, we are RECONCILING -- go to "Reconcile". - If none exists, we are SETTING UP -- go to "Set up". SET UP (first time): 3. Ask me for the prompt and the model(s) to monitor. List the models available to my account first and respect my plan's model-per-comparison cap -- do not exceed it. 4. Make sure the benchmark is SCORED, or drift in quality cannot be detected: ask me for a gold-standard answer (expected_output) or a saved rubric (rubric_id). If I have neither, offer to help me build a rubric first. 5. Create a scheduled benchmark (create_benchmark) with the prompt, models, scoring, and a schedule I choose (daily / 4x_daily / hourly -- tell me which my plan allows; the server enforces the cap). Note the suite_id it returns. 6. Optionally set alert thresholds so regressions are flagged: ask me what I care about ("warn if cost jumps 20% or quality drops 10%") and translate it with update_benchmark_alerts. These thresholds drive the anomaly_flags you will read later. 7. Record the standing note as the series anchor: put_agent_note(recipe_id="recipe_id", recipe="mcp-drift-monitor", intent="Drift monitor for <model(s)> on <short prompt label>", status="running", data={ suite_id: "<the suite_id>" }). Keep status "running" -- this note is the anchor for the whole series, not a one-shot. 8. Trigger a first run now (trigger_benchmark_run) to establish the baseline; when it returns a job_id, update the note's job_id (keep status "running"). Poll it to completion (poll_job every 5 seconds, stop after 5 minutes), then tell me the baseline numbers per model. Leave the note "running". RECONCILE (every later session, or when I ask "any drift?"): 3. From the standing note, read the suite_id (note.data.suite_id). 4. Pull the recent runs: list_benchmark_runs(suite_id). Each run carries anomaly_flags -- the server-computed regression flags vs the immediately-prior run (cost_spike, latency_regression, score_drop, model_version_change, threshold_breach). Read the newest run's anomaly_flags FIRST -- do not recompute deltas yourself. 5. If I want to compare against a specific earlier baseline (e.g. "vs last week", not just the prior run), call diff_runs(baseline_run_id=<the earlier run>, current_run_id=<the newest run>). It returns per-model deltas (prev, curr, delta, pct for cost / latency / quality / judge) plus regression flags -- again, server-computed, so never subtract numbers in your head. 6. For the shape of the trend over many runs, call get_benchmark_trend(suite_id) and read the per-model run_series. 7. Report in plain language, and BE QUIET WHEN NOTHING MOVED: if there are no flags and no material deltas, say "no drift since the last check" in one line. Only when something actually moved do you lay out which model, which metric, by how much, and since when. 8. Re-point the standing note for the next cycle: update the note's job_id to the newest run and keep status "running". Do NOT mark it completed -- a standing series stays "running" as its anchor. ENDING THE SERIES (only when I explicitly say to stop monitoring): - Mark the standing note "completed" (it drops out of the active list and frees a slot). Optionally delete the benchmark if I no longer want it to run. ON ANY FAILURE or clear mismatch (at any step): tell me honestly -- do not fabricate a result or a "no drift" all-clear. Mark the standing note "failed" and stamp its data with failure telemetry: failure_stage, failure_reason, a one-line detail, and the job_id if a run was involved. Never leave the note "running" after a genuine failure, and never report "no drift" when you could not actually read the runs.
Goal
Turn a one-off “which model is best?” decision into ongoing assurance that your choice stays best. You stand up a scored benchmark on a schedule once; from then on your agent reconciles each run against the baseline and tells you – in plain language, only when it matters – if quality dropped, cost spiked, or latency crept up. Nobody watches a chart.
When to use
Reach for this after you have chosen a model or pipeline for a real task and you care whether it holds up: provider-side model updates, silent quality regressions, cost drift. It is the standing-series sibling of a one-off comparison – the same measurement, running on a cadence, reconciled across sessions so the signal finds you instead of the other way around.
How it works
- One standing note is the anchor. Unlike a one-shot run (which the agent marks
completedwhen done), a drift monitor keeps its noterunningfor the life of the series. Each cycle re-points the note at the newest run; the note is only closed when you end monitoring. That is what lets any future session pick the series up by itsrecipe_id. - The server does the arithmetic. Drift detection leans on three server-computed
signals so the agent never eyeballs numbers: each run’s
anomaly_flags(regression vs the prior run),diff_runs(per-model deltas vs any baseline you choose), andget_benchmark_trend(the time-ordered series). The agent interprets and advises; the classical layer finds the pattern the same way every time. - Silence is a feature. When nothing moved, the agent says so in one line. The report is loud only when there is real movement to act on.
Verification
- A scheduled benchmark exists and has run at least once (the baseline).
- A standing note exists with the
suite_idin itsdata, statusrunning. - Asking “any drift?” in a fresh session finds the note, reads the latest run, and reports movement (or a one-line all-clear) without you supplying a job_id.
- The note stays
runningacross cycles and only becomescompletedwhen you end the series; a genuine failure leaves itfailedwith afailure_reason, never a false “no drift”.