MCP Recipes
ProSet Up Alerts on a Benchmark
Describe what you want to be warned about in plain language -- a cost spike, a quality drop, a latency jump -- and the agent translates it into alert thresholds on a benchmark.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, turn a plain-language alerting intent into alert thresholds
on a benchmark. This is configuration only -- you are not running anything that costs money.
1. Confirm the tools you need are available (list the tools).
2. Identify the benchmark: if I did not name one, list_benchmarks and ask me which to set
alerts on.
3. Read the current alert config first: get_benchmark_alerts(suite_id). Tell me what is
already set so we are changing from a known baseline, not blind.
4. Ask me what I want to be warned about if I have not said it, and translate it into the
specific thresholds:
- cost_spike_pct: alert when a run's cost rises this % vs the previous run
- latency_regression_pct: alert when latency rises this % vs the previous run
- quality_drop_pct: alert when the quality score drops this % vs the previous run
- threshold_breach_alert: alert when any model fails its configured pass/fail threshold
Map my words to numbers explicitly and show me the mapping before applying (e.g. "'big
cost jump' -> cost_spike_pct: 25"). Set a metric to null to leave it unalerted; do not
invent thresholds I did not ask for.
5. Apply it: update_benchmark_alerts(suite_id, enabled=true, <the thresholds>, [email]).
If I did not give an email, it defaults to my account email -- tell me that.
6. Read the config back and confirm in plain language exactly what WILL and WILL NOT alert,
so there are no surprises. Note that webhook delivery is Enterprise-only; if I asked for
a webhook and am not on Enterprise, say so rather than silently dropping it.
ON ANY FAILURE (benchmark not found, update rejected): tell me honestly what happened --
do not claim alerts are set if the update did not succeed.
Goal
Get useful alerts configured without learning the threshold fields. You say what you care about in plain language; the agent maps it to the right thresholds on your benchmark, shows you the mapping, applies it, and confirms exactly what will and won’t trigger.
When to use
Reach for this once you have a benchmark you rely on and want to be told when it moves – cost, latency, quality, or a threshold breach – instead of watching it. It is pure configuration: nothing is run, nothing is charged.
How it works
- Plain language in, thresholds out. The agent maps intent (“warn on a big cost jump”)
to the specific fields (
cost_spike_pctetc.) and shows the mapping before applying, so you approve the numbers. - Baseline-aware. It reads the current config first and reads it back after, so you see the before and after and exactly what will and won’t alert.
- Honest about tier limits. Webhook delivery is Enterprise-only; the agent says so rather than silently dropping a webhook request.
Verification
- The current alert config was read before changes and read back after.
- Every threshold set maps to something you asked for; metrics you didn’t mention are left unalerted (null), not invented.
- The final confirmation states plainly what will and will not trigger an alert.
Where this leads
Natural next recipes once you are comfortable with this one.