Skip to content

Close your laptop. Your agent remembers.

· 7 min read

Some jobs do not finish in one session. A drift monitor watches a benchmark across weeks. An optimization loop runs through an iteration budget that takes longer than a chat window. A data pipeline kicks off a long scoring run and the user closes the laptop before it completes.

These all need the same thing: a way for a fresh session to find out what the previous session was doing and pick it up.

LLM Prover’s MCP server provides exactly that. It gives an agent persistent memory through a set of note tools, so a long-running workflow, loop, or graph can survive the end of a session and resume seamlessly in the next one. The agent writes down where it is; a later session reads it back and continues. That persistence is what turns a one-shot chat into a standing job.

How the agent remembers

The memory is a durable note. The agent writes a small, structured record to the LLM Prover server using put_agent_note, and that record lives there until the agent marks it complete or deletes it. Every note is tagged with a recipe_id so a later session reads it back by asking for notes under that ID.

A drift-monitor note holds the benchmark it is watching, the baseline run, and the last run that was reconciled. An optimization-loop note holds the iteration count, the current best prompt, the score, and the spend. The note is not the result. It is a bookmark: the minimum state the next session needs to continue.

A worked example: a drift monitor

To make this concrete, the rest of this post follows one example end to end: a drift monitor. A drift monitor is a standing job that re-runs a benchmark on a schedule and flags when a model’s quality, cost, or latency moves. It is a natural fit for persistent memory, because the whole point is to keep watching across days and weeks, long past any single session.

The monitor set up for this post watches one benchmark: a support-ticket summarisation task scored across seven models on a daily schedule.

What the monitor watches

Benchmark: Customer support tickets

Prompt: Summarise the support ticket in exactly three bullet points.

Models tracked: Grok 4.5, Claude Opus 5, GPT-4o Mini, GPT-4.1 Nano, DeepSeek V4 Flash, Kimi K3, Qwen 3.8 27B

Schedule: daily

With the target defined, the agent anchors the monitor:

Agent
> create_benchmark (support-ticket task, daily, rubric)
benchmark created, 7 models
> trigger_benchmark_run
baseline run complete
> put_agent_note (recipe_id, benchmark, baseline_run, status: running)
note stored | status: running | baseline anchored

Here is the note that call wrote, as it comes back from the server:

{
  "recipe_id": "drift-monitor-001",
  "recipe": "mcp-drift-monitor",
  "status": "running",
  "intent": "Drift monitor for support-ticket task across 7 models",
  "data": {
    "benchmark": "Customer support tickets",
    "baseline_run_id": "813749a0...",
    "last_run_id": "813749a0..."
  }
}

Small on purpose. It is a bookmark, not a result store, holding just:

  • which benchmark it watches
  • which run is the baseline
  • which run was reconciled last
  • why the note exists

The whole state a future session needs fits in a handful of fields.

The key detail is the status. It stays running. Unlike a one-shot run, a standing monitor does not complete after one cycle. The note stays active as the anchor for the whole series, its last_run_id updated each time the agent reconciles, and it is only marked completed when the user ends the monitoring.

A fresh session picks it up

The “Reconcile on Startup” recipe sets this up. With your approval, it adds a standing instruction to your agent’s own context, the same way a startup instruction or project rule persists across sessions, so every new session begins by checking its notes before you ask. The recipe shows you the exact text and saves nothing until you say yes. Here is the instruction it proposes:

The standing instruction (saved to your client, with approval)

On session start, call list_watched_recipes and show me any active (not completed) watched recipes as a short list, name and what each is. Ask me which, if any, I want to resume now; if I say none (or want to do something else), do nothing further with LLM Prover. For each one I choose, read its note (get_agent_notes) and resume it where it left off according to its own type. Never auto-resume, poll, or spend without my selection.

Notice that this default is conservative: it lists what is active and waits for you to choose, and spends nothing without your say-so. That is a choice, not a constraint. The instruction is your text, in your context, so you tune it to your workflow. A cautious operator keeps the human checkpoint. A team running LLM Prover as a node inside an agentic graph can drop it, instructing the agent to resume and reconcile every active note automatically, no approval, because the whole pipeline is meant to run unattended.

Insight: The resume behaviour is not fixed by the product. It is an instruction you own and edit. Keep a human in the loop when you want control, or remove the checkpoint and let the agent resume everything on its own when the pipeline is designed to run without you. Same memory, different autonomy.

With that instruction in place, the behaviour below needs no prompting. It was tested cold: a new session reached for list_watched_recipes, found the active note with no ID supplied by the user, and read back the full state shown above.

Agent
> list_watched_recipes
1 active note: mcp-drift-monitor (no id supplied by user)
> get_agent_notes (recipe_id)
note found | benchmark + baseline_run in state
> list_benchmark_runs
newest run differs from last reconciled
> diff_runs (baseline, newest)
flags: latency_regression grok-4.5 +53% | cost_spike deepseek +100%
> update_agent_note (last_run_id=newest, status=running)
note updated | 2 flags carried to report

No human handed the agent a benchmark ID. No one told it what it was watching. It discovered the note, read the state, pulled the latest run, and diffed it against the baseline it had saved. The diff is server-computed, so the agent reports the flags rather than eyeballing the numbers. In this reconcile, two models had moved since the baseline: Grok 4.5 latency rose 53% and DeepSeek V4 Flash cost doubled. The agent reported both and re-pointed the note at the newest run for next time.

When nothing has moved, the same path reports a one-line all-clear. Loud when there is something to act on, quiet when there is not.

That is the payoff: persistent memory plus a startup instruction, picking up a baseline set in a session that had already ended and reporting the exact two models that had drifted, work a human would otherwise redo from scratch each time.

Run jobs that outlast a single session

Agent-state and the full recipe catalog are on Pro and above.

Get started on Pro

What the state layer handles

The note tools are deliberately simple:

  • put_agent_note creates or replaces a note, scoped to a recipe ID
  • get_agent_notes reads notes for a recipe, returning active ones by default
  • update_agent_note re-points the job ID, updates data fields, or changes status
  • list_watched_recipes shows all active notes across recipes (the “what is the agent watching” view)

Active notes are capped per tier (the server enforces the limit). Completing a note frees the slot. Standing series stay running as their anchor; one-shot jobs complete when done and open the slot for the next job.

Where this shows up in the recipes

Four of the shipped recipes use agent-state:

  • Drift monitor: standing note anchors the monitoring series across sessions
  • Optimization loop: note holds iteration count, best prompt, score, and spend
  • Track a run: note records a long job so a fresh session can reconcile on startup
  • Reconcile on startup: checks all active notes and reports anything that finished

Each one writes the note after every step and reads it back to confirm the state landed. The note is the source of truth, not the agent’s working memory. If the session dies mid-step, the next session reads the last confirmed state and picks up from there rather than re-running.

What’s next

mcp agent-state agentic benchmarking