MCP Recipes
ProTest Your Chatbot With a Conversation Scenario
Declare a scenario -- a persona, a goal, what to probe, and what good looks like -- and have the agent hold a grounded multi-turn conversation with your chatbot, scoring each turn so you see exactly where it does its job and where it slips.
Agentic and programmatic use. Agents and scripts you connect can make cost-generating calls on your behalf, and AI agents behave non-deterministically. You are responsible for monitoring and bounding your own automation. See Terms for details.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, test a CHATBOT across a multi-turn conversation. This is declarative and MCP-native: I give you a SCENARIO (not a fixed script); you drive the conversation to satisfy it, reacting to what the bot actually says, and score each turn. A conversation is ONE stateful thread for ONE model -- you never run several models in one thread (after turn 1 they diverge and can't be compared). Keep status "running" until the conversation ends. SET UP: 1. Confirm the tools you need are available (list the tools): run_evaluation, poll_job, list_models, list_rag_stores, create_rubric, list_rubrics, delete_rubric, get_agent_notes, put_agent_note, update_agent_note. (You score each turn with a saved rubric built from the scenario's probes -- run_evaluation has no "inline criteria" parameter, so the rubric is how the probes reach the judge; see step 6b.) 2. Read your notes: get_agent_notes(recipe_id="recipe_id"). If an active note exists we are RESUMING -- read its data (scenario, grounding, rubric_id, turn, cost_spent, per-turn scores+evaluation_ids, last job_id, terminate/outcome); reuse the stored rubric_id for scoring (do not create a second rubric); rebuild the transcript from the stored run records' verbatim response_text (step 8 fidelity rule), not from memory; and FIRST re-check the stop contract against the counters (step 7); if a ceiling or the goal was already reached, go to REPORT, do not keep going. If no active note, we are STARTING. 3. GET THE BOT UNDER TEST. Call list_models and have me pick the ONE model that is my chatbot (or tell you my BYOE endpoint). Never hardcode a model id. One model only. 4. GET THE GROUNDING -- MANDATORY. A bare model with no grounding is NOT my chatbot; testing it measures the model's general knowledge, not my bot. Require at least ONE of: - system_prompt_id (my bot's instructions/persona), - store_id (my bot's knowledge base), or - context (inline policy/doc text), (a BYOE endpoint may already carry its own grounding -- that counts). If I give NONE of these, STOP and tell me plainly that the test would be meaningless without the grounding my real bot uses. If grounding is a store_id, VERIFY it first: list_rag_stores -- it MUST be sync_status "ready" (if "syncing", wait; if "dirty", REFUSE -- a conversation grounded in a broken store scores the store, not the bot, at real cost). Suggest proving retrieval first (the "Prove a RAG Store Actually Retrieves" recipe) if unsure. 5. GET THE SCENARIO from me (declarative -- what the conversation should achieve, not the user's exact words): - persona (who the simulated user is), - goal (what they are trying to accomplish), - must_probe (the behaviours to test, e.g. "asks for an order number before acting", "stays on refund policy for a >14-day-late order", "stays calm if the user escalates"), - success (what a good outcome looks like), - guardrails (what the bot must NOT do, e.g. promise a refund it cannot authorise), - and the STOP CEILINGS: a max turn count and a cost cap. The must_probe list IS the scoring rubric -- you score each turn against these. 5b. SETTLE THE SCORING RUBRIC with me -- do NOT silently autogenerate it. The rubric is the scoring instrument for the WHOLE run; a quietly-built one buries a decision I should own, and (as the probe caveat below shows) its phrasing can swing the whole trajectory. The scoring vehicle is a saved rubric: run_evaluation accepts expected_output, rubric_id, or rubric_ids to score against, but has NO free-text "inline criteria" input, and its context field is RAG grounding, not a rubric -- so the probes must become a rubric_id one way or another. First call list_rubrics so you can offer my existing ones, then ASK me which of these four I want, and wait for my answer: (1) USE AN EXISTING RUBRIC -- I name one from list_rubrics; use its rubric_id as-is. (2) CO-CREATE IT WITH ME -- walk the must_probe list with me criterion by criterion (name, weight, description), agree each, then create_rubric from what we settled. (3) PROPOSE ONE FOR MY APPROVAL -- you draft criteria from must_probe and show them to me as text; create_rubric only AFTER I approve or amend. (4) JUST GENERATE ONE AND CONTINUE -- you build it from must_probe without a checkpoint (fastest; use only if I say so). For any path that creates a rubric (2/3/4): create_rubric(name="<scenario> probes", rubric_type="quality", criteria=[{criterion, weight, description} per probe]); give each a weight and a SPECIFIC description (vague descriptions score inconsistently). Rubric criteria count is tier-capped -- keep must_probe within my plan's cap; create_rubric reports the limit if exceeded. Keep the returned rubric_id; every turn's run_evaluation passes it, and it goes in the note so a resume reuses it. WRITE PROBES QUERY-CONDITIONAL (applies to options 2/3/4, and flag it if an existing rubric from option 1 has the problem): a probe phrased as an unconditional demand (e.g. "always gives actionable next steps") will score 0 on a turn whose user question does not call for it (a scope or yes/no availability question), dragging the turn score down even when the bot answered correctly. Phrase such probes conditionally ("gives actionable next steps WHEN the user asks how to do something") or read the trajectory knowing a dip may be a probe-fit artifact, not a bot failure -- confirm at the human checkpoint (step 13). 6. Record the note BEFORE spending, with the scenario, grounding, and the structured stop contract so a later session reads its own limits: put_agent_note(recipe_id="recipe_id", recipe="mcp-chatbot-scenario-test", status="running", intent="Chatbot scenario test: <goal> for <model>", data={ model|byoe, grounding:{system_prompt_id|store_id|context}, scenario:{persona, goal, must_probe[], success, guardrails}, rubric_id:<from step 5b>, turn:0, cost_spent:0, scores:[], last_job_id:null, terminate:{ iterations:<max turns>, cost_usd:<cost cap> }, outcome:{ goal_resolved:false } }). NOTE: the note is a POINTER, not a transcript store (it has a ~16KB cap). It holds turn count, per-turn scores, the scenario, and each turn's evaluation_id + last job_id. The AUTHORITATIVE transcript is the sequence of run records, re-fetchable by those evaluation_ids -- NOT the agent's working memory, which is a convenience copy only. A fresh session rebuilds the real conversation by reading the stored run records' verbatim response_text in turn order (step 8 fidelity rule), so the transcript is auditable and the agent cannot silently rewrite history. Never try to store the whole transcript in the note. LOOP (each turn): 7. CHECK THE STOP CONTRACT against the note's counters -- OUTCOME first, then TERMINATE; first to trip wins: a. OUTCOME: if the scenario goal is resolved (success criteria met) -> STOP as SUCCESS; REPORT then mark the note completed with stop_outcome "goal_resolved". b. TERMINATE: if turn >= terminate.iterations OR cost_spent >= terminate.cost_usd -> STOP as CEILING; REPORT then mark completed with stop_outcome "halted_at_ceiling" and stop_detail (which fired). A ceiling hit is TERMINAL even with no human present -- never a zombie that keeps chatting. 8. GENERATE THE NEXT USER TURN to advance the scenario -- a realistic message this persona would send next, INFORMED by what the bot actually said last turn (this is the point of declarative: adapt, don't replay a blind script). On turn 1 it is the opening message. FIDELITY RULE (non-negotiable -- this is what makes the test real): YOU generate the USER turns, but the BOT's prior turns in the transcript MUST be the bot's own words, verbatim, from the stored run records -- NEVER your paraphrase, summary, trim, or "tidied" version. You are not the system of record for what the bot said; the server is. Before building this turn's prompt, retrieve the prior turn's actual output: read response_text from the last run (it is in the poll_job result when complete, and re-fetchable via the stored evaluation_id / list_evaluations). Paste that verbatim as "Assistant (turn N): ...". If you ever cannot retrieve a prior turn's real response_text, STOP and mark the note failed (reason: transcript_unverifiable) -- do NOT reconstruct it from memory. Silently smoothing a hedge or "fixing" a contradiction you noticed would measure a conversation that never happened and destroy exactly the self-referential signal (does the bot stay consistent with what it ACTUALLY said earlier?) this recipe exists to catch. 9. SEND IT TO THE BOT, grounded and in context: run_evaluation( prompt=<the full conversation so far, formatted as a transcript: system/grounding + all prior turns, where each prior USER turn is your generated message and each prior BOT turn is that turn's VERBATIM response_text from its run record (step 8 fidelity rule) + this new user turn>, models=<the one model>, system_prompt_id=<if that is the grounding>, store_id=<if that is the grounding>, context=<if that is the grounding>, rubric_id=<the rubric built in step 5b from must_probe>). Async -- poll_job every 5 seconds to completion. The scored "response" is the bot's next turn; read results[].quality_score, the per-criterion breakdown/judge_reasoning, the evaluation_id, and the VERBATIM response_text (that text is turn N's authoritative bot reply -- carry it forward into the next turn's prompt unchanged). (If I gave a BYOE endpoint, target it as the model/provider instead.) 10. RECORD: append {turn, user_turn, evaluation_id, bot_response_summary, score, per_probe_scores, cost} to the note's scores[] -- store the evaluation_id so the verbatim bot reply is always re-fetchable and the transcript stays auditable; bot_response_summary is a human-readable convenience ONLY, never the text you feed back (that is always the run record's response_text, per step 8). Set turn += 1; cost_spent += this turn's cost; last_job_id = this job. READ THE NOTE BACK (get_agent_notes) and confirm the write took. The note is the source of truth for the trajectory; the run records are the source of truth for the transcript; your working memory is neither. 11. Go back to step 7 (re-check stop contract, then next turn). REPORT (when a stop condition fired): 12. Give me the PER-TURN SCORE TRAJECTORY (score per turn, so I can see where the bot held up and where it slipped), WHICH stop condition ended it (goal resolved vs ceiling), and a CROSS-TURN INSIGHT -- not just numbers: where did it do its job, where did it degrade, did it breach a guardrail, did it lose the thread as history grew? Name the turn where quality dropped if it did. Quote the bot's own words for the key moments. 13. HUMAN CHECKPOINT (overfit caution): a per-turn score is a proxy, and scoring to a judge can be gamed -- a bot can earn high scores with confident, verbose, judge-pleasing replies that are not actually correct or policy-consistent (overfit-to-the-judge / Goodhart). Do NOT trust the score alone. Where ground truth exists, check against it -- a store-grounded answer against the source document, a policy probe against the actual policy text -- not just the judge's opinion. Show me the actual conversation (or the pivotal turns), rebuilt from the run records' verbatim response_text (not your memory), not just the trajectory, and recommend I read it before trusting the verdict. RUN-TO-RUN VARIANCE: the bot is stochastic, so the same scenario can score differently across runs -- a single run is a sample, not a verdict. Do not read one run's dip as a regression; to track real change over time, lock a scenario that passes and re-run it on a cadence (the drift-monitor recipe). Flag to me if a result looks like noise rather than signal. 14. CLEANUP: mark the note "completed" with stop_outcome (and stop_detail for a ceiling). The scoring rubric from step 5b is scenario-specific scaffolding -- if it was created only for this run, offer to delete_rubric it (keep it if I want to reuse the same probes to compare models). Never delete my grounding store. To compare MODELS, tell me to re-run this same scenario against another model and compare trajectories -- reuse the same rubric_id so the probes are identical across runs; one model per conversation, so comparison is across runs, not within one. ON ANY FAILURE or clear mismatch (a turn will not score, grounding store went not-ready mid-run, the note write did not read back, you lost the conversation thread): STOP, mark the note "failed", and stamp data with failure_stage, failure_reason, a one-line detail, and the last good job_id. Never keep spending on a conversation whose state you cannot trust, and never report a verdict you did not actually measure.
Goal
Prove your chatbot does its job across a real conversation, with data – not a single prompt. You declare a scenario (a persona, a goal, what to probe, what good looks like); the agent holds a grounded multi-turn conversation with your bot, generating each user turn to advance the scenario and reacting to what the bot actually says, and scores every turn. You get a per-turn score trajectory that shows exactly where the bot holds up and where it slips – the thing single-prompt testing cannot see.
When to use
Reach for this when you run a chatbot (support, onboarding, sales) and “is it any good?” needs an evidence-based answer. Especially before shipping a change to the bot’s prompt or knowledge base, or to compare candidate models on the same conversational job.
How it works, and its limits (read this)
- Declarative, not scripted. You declare what the conversation should achieve and probe, not the user’s exact words. The agent drives the conversation to satisfy the scenario, adapting each user turn to the bot’s real responses. This is the MCP-native shape: the agent is the test driver, not a replayer of a brittle fixed transcript. (A strict literal script – exact user turns – is possible for pure regression, but it is blind to the bot’s replies and reads oddly when the bot answers unexpectedly; declarative is the better default.)
- One model per conversation. A conversation is a single stateful thread – each turn depends on that model’s own previous replies. You cannot run several models in one thread (after turn 1 they diverge). To compare models, run the SAME scenario against each separately and compare the trajectories. The reproducible axis is the scenario (persona + goal + probes), not identical words.
- Grounding is mandatory. A bare model with no system prompt, store, or context is not your
chatbot – testing it measures the model’s general knowledge. The recipe requires grounding
(system prompt / RAG store / pasted context / a BYOE endpoint that carries its own) and
refuses to run without it. If grounded by a store, it checks the store is
sync_status: readyfirst – a conversation grounded in a broken store scores the store, not the bot, at real cost. - Scored per cumulative turn. Each turn is scored given everything before it, against the
scenario’s
must_probecriteria. The output is a trajectory – so a bot that is great for three turns then loses the thread shows up as a score that climbs then drops, localising the failure. Context retention and consistency are exactly what this measures. The probes become a saved rubric, and because that rubric is the scoring instrument for the whole run, the agent settles it with you up front rather than inventing one silently – use a rubric you already have, co-create it, approve a proposed one, or tell it to just generate one. One caveat on reading the trajectory: a probe written as an unconditional demand (“always gives actionable next steps”) will score a turn down when the user’s question did not call for it (a scope or yes/no availability question) – the bot can be entirely correct and still dip. Phrase probes query-conditionally, or treat a dip on a turn whose question did not invoke that probe as a probe-fit artifact and confirm it at the human checkpoint before reading it as a bot failure. - The transcript is the bot’s real words, not the agent’s. The agent writes the USER turns,
but every prior BOT turn fed back into the conversation is that turn’s verbatim
response_textpulled from its stored run record (kept byevaluation_id), never a paraphrase, trim, or “tidied” version. This matters most for self-referential probes (“earlier you said X – are you sure?”): if the agent could quietly smooth a hedge or fix a contradiction it noticed, the test would measure a conversation that never happened and the consistency signal would vanish. The run records are the authoritative transcript, so a run is auditable after the fact and the agent has no discretion to rewrite history; if a prior turn’s real text can’t be retrieved the run fails rather than reconstructs from memory. (Note: the agent still hand-assembles the transcript string from those records – the one thing it cannot do is alter the bot’s words. Server-side threading, where you pass prior run ids and the server assembles history, is a stronger guarantee and is on the roadmap, not in this recipe yet.) - Non-determinism is expected. The bot is stochastic: the same scenario can score differently run to run. A single run is a sample, not a verdict – don’t read one run’s dip as a regression. To track genuine change over time, lock a scenario that passes and re-run it on a cadence (the drift-monitor recipe); that is what separates “the bot varied” from “the bot got worse”.
- Bounded and resumable. The loop carries a
terminateceiling (max turns + cost cap) and anoutcome(goal resolved) in the note, checked each turn and on resume; it stops on the first to trip and never runs away. The note holds a pointer (turn count, per-turn scores + evaluation_ids, scenario, last run id) – not the full transcript; a resume rebuilds the real conversation from the stored run records, not from memory. - The judge is a proxy – watch for overfitting it. A high per-turn score is not proof the bot is genuinely good. Scoring to a judge can reward judge-pleasing replies (confident, verbose, well-formatted) over actually-correct, policy-consistent ones – the usual overfit-to-the-judge / Goodhart risk. Where ground truth exists, prefer it over the judge’s opinion: a store-grounded answer can be checked against the source document, and a policy-consistency probe against the actual policy text, not just “did it sound right”. The recipe ends with a human checkpoint – read the actual conversation, not just the numbers, before trusting the verdict.
Verification
- Exactly one model drove the conversation, and it was grounded (system prompt / store /
context / BYOE) – an ungrounded run was refused. A store grounding was
sync_status: ready. - Each turn was scored against the declared
must_probecriteria, with the conversation carried forward so later turns were judged in full context. - The report is a per-turn trajectory plus a cross-turn insight (where it held up, where it slipped, any guardrail breach), not a single number – and names the turn where quality dropped if it did.
- The loop had a
terminateceiling set before any spend and stopped on the first condition to trip; a ceiling hit ended terminally. A human checkpoint on the actual conversation was offered before the verdict. - To compare models, the same scenario was re-run per model and trajectories compared – never multiple models in one conversation.
Where this leads
Natural next recipes once you are comfortable with this one.
If you ground the bot with a RAG store, first prove the store actually retrieves the right content -- a conversation grounded in a bad store scores the store, not the bot.
Lock in a scenario that passes, then re-run it on a cadence to catch the day your bot's conversation quality drifts.