MCP Recipes
pro+ onlyTrack a Run Across Sessions
Kick off a long-running LLM Prover job, remember it with an agent-state note, and reconcile it in a later session instead of losing track.
Run it -- paste into your MCP agent
Highlighted parts are placeholders -- replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, start a job and track it durably so we can pick it up later: 1. Confirm the tools you need are available (list the tools). 2. Record a note BEFORE starting work, so a crash between launch and note cannot orphan the run: put_agent_note(recipe_id="recipe_id", recipe="mcp-track-a-run", intent="what you are running and why", status="running"). 3. Start the async operation (run_comparison, run_evaluation, or trigger_benchmark_run). When it returns a job_id, update the note with it: update_agent_note on this note, setting job_id and keeping status "running". 4. Poll poll_job every 5 seconds until it is complete or failed; stop after 5 minutes. (If I close the session, a later session can read this note and resume polling.) 5. Report the result to me in your own words. 6. Cleanup: update the note to status "completed" (it then drops out of the active list and frees an active-note slot). 7. On ANY failure or clear mismatch: tell me honestly (do not fabricate a result), mark the note "failed", and stamp the note's data with failure telemetry: failure_stage, failure_reason, a one-line detail, and the job_id if there was one. Never leave the note "running" after a failure; never mark a failed run "completed".
Goal
Make a long-running job survive the session. You start the job, the agent records a durable note tracking it, and even if this session ends, a later session can find the note, poll the job, and tell you the result – instead of you having to remember a job_id.
When to use
Any time a job is slow enough that you might not want to sit and watch it, or when you want an agent to kick things off and report back later. This is the foundational pattern behind every multi-session recipe (weekly drift checks, scheduled benchmarks, long comparisons): record a note, reconcile it next time.
Steps
- Record first. The agent writes an agent-state note with status
runningand your intent, before starting the job – so nothing is orphaned if the launch and the note don’t both land. - Start and link. The job returns a
job_id; the agent stores it on the note. - Poll. The agent polls to completion (or hands off to a later session, which reads the note and resumes).
- Report and clean up. On success, the agent reports the result and marks the note
completed– which frees an active-note slot. - Fail honestly. On failure, the agent marks the note
failedand records structured telemetry (stage, reason, detail) so failures become measurable signal.
Verification
- A note exists for this run with the correct
job_id. - On success, the note ends
completedand you got the result. - On failure, the note ends
failedwith afailure_reason– never leftrunning. - Starting a fresh session and asking “what am I waiting on?” surfaces this run while it is still active, and nothing once it is completed.
Where this leads
Natural next recipes once you are comfortable with this one.