Skip to content

MCP Recipes

Pro

Review Your Recent Usage and Spend

Have the agent pull your recent comparisons, evaluations, and benchmark runs and summarise cost, latency, and activity -- a quick read on what you've been running and what it's costing.

Run it: paste into your MCP agent

Highlighted parts are placeholders. Replace them with your own values before (or after) copying.

Using the LLM Prover MCP tools, give me a read on my recent usage and spend. This is
read-only -- do not create, run, or delete anything.

1. Confirm the tools you need are available (list the tools).
2. Pull my recent activity: list_comparisons and list_evaluations (recent first). If I
   have benchmarks, also list_benchmarks and, for the ones I care about, list_benchmark_runs.
3. Summarise in plain language: how much I've run lately across comparisons / evaluations /
   benchmarks, the rough total and per-item cost where it is available, and any latency or
   cost outliers worth noticing. Group by activity type so it is easy to scan.
4. Point out anything actionable: an unusually expensive run, a model that is consistently
   slow or costly, a benchmark that has not run recently. Keep it factual -- this is a
   snapshot, not advice to act on blindly.
5. If I re-run the same task on a cadence, mention (once) that a drift monitor could watch
   it for me instead of manual checks -- only if it genuinely fits what you see.

ON ANY FAILURE (a list call errors, nothing comes back): tell me plainly what you could and
could not retrieve -- do not fabricate usage numbers.

Goal

A fast, read-only snapshot of what you’ve been running and what it’s costing – comparisons, evaluations, and benchmark runs summarised in plain language, with the expensive or slow items flagged, so you don’t have to page through dashboard history.

When to use

Reach for this when you want a quick sense of recent activity and spend, or before deciding whether a recurring task is worth automating. It touches nothing – it only reads.

How it works

  • Read-only. The recipe only lists existing records (comparisons, evaluations, benchmark runs). It never creates, runs, or deletes anything.
  • Summary, not a dump. The agent groups by activity type and surfaces cost/latency context and outliers, rather than echoing raw rows.

Verification

  • The summary reflects your actual recent records (comparisons / evaluations / benchmark runs).
  • Cost and latency figures come from the records, not invented.
  • Nothing was created, run, or deleted.

Where this leads

Natural next recipes once you are comfortable with this one.