MCP Recipes
ProReview Your Recent Usage and Spend
Have the agent pull your recent comparisons, evaluations, and benchmark runs and summarise cost, latency, and activity -- a quick read on what you've been running and what it's costing.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, give me a read on my recent usage and spend. This is read-only -- do not create, run, or delete anything. 1. Confirm the tools you need are available (list the tools). 2. Pull my recent activity: list_comparisons and list_evaluations (recent first). If I have benchmarks, also list_benchmarks and, for the ones I care about, list_benchmark_runs. 3. Summarise in plain language: how much I've run lately across comparisons / evaluations / benchmarks, the rough total and per-item cost where it is available, and any latency or cost outliers worth noticing. Group by activity type so it is easy to scan. 4. Point out anything actionable: an unusually expensive run, a model that is consistently slow or costly, a benchmark that has not run recently. Keep it factual -- this is a snapshot, not advice to act on blindly. 5. If I re-run the same task on a cadence, mention (once) that a drift monitor could watch it for me instead of manual checks -- only if it genuinely fits what you see. ON ANY FAILURE (a list call errors, nothing comes back): tell me plainly what you could and could not retrieve -- do not fabricate usage numbers.
Goal
A fast, read-only snapshot of what you’ve been running and what it’s costing – comparisons, evaluations, and benchmark runs summarised in plain language, with the expensive or slow items flagged, so you don’t have to page through dashboard history.
When to use
Reach for this when you want a quick sense of recent activity and spend, or before deciding whether a recurring task is worth automating. It touches nothing – it only reads.
How it works
- Read-only. The recipe only lists existing records (comparisons, evaluations, benchmark runs). It never creates, runs, or deletes anything.
- Summary, not a dump. The agent groups by activity type and surfaces cost/latency context and outliers, rather than echoing raw rows.
Verification
- The summary reflects your actual recent records (comparisons / evaluations / benchmark runs).
- Cost and latency figures come from the records, not invented.
- Nothing was created, run, or deleted.
Where this leads
Natural next recipes once you are comfortable with this one.