Skip to content

Score every turn of a chat, not just the first reply

· 9 min read

You run a chatbot. Maybe it onboards new customers, maybe it answers support questions, maybe it helps a prospect decide to buy. You tested it the obvious way: you typed in a hard question, read the reply, and it was good. So you shipped it.

The problem is that your customers do not send one message. They have a conversation. They ask a follow-up, they refer back to something the bot said three turns ago, they wander off topic and come back. A single-prompt test proves the bot can produce one good answer. It says nothing about whether the bot holds up across a real exchange, where the hard part is memory, consistency, and knowing what it does not know.

A new LLM Prover recipe tests the thing you actually ship: the conversation. You hand your agent a scenario, it holds a grounded multi-turn chat with your bot, and it scores every single turn. The output is a trajectory, a score for each turn, so you can see the exact point where a bot that started strong begins to slip.

The setup, before any scores

To make this concrete, here is a real test run against a real bot: an onboarding assistant for new staff, grounded in an actual knowledge base about a product. Everything below is the input to the test. The scores come after.

The bot under test is one model with grounding. A bare model with no grounding is not your chatbot, it is a general-knowledge model, so the recipe refuses to run without it.

The chatbot under test

  • Model: gpt-4o-mini (a cheap, common choice for a production support bot)
  • Grounding: a RAG store holding the product knowledge base (features, tiers, pricing)
  • System prompt: “You are the onboarding assistant. Answer using only the provided knowledge base. If something is not covered, say you do not know rather than guessing. Never invent a feature, price, or policy.”

The scenario is declarative. You do not script the exact words. You declare what the conversation should achieve and what to probe, and the agent generates each user turn to fit, reacting to what the bot actually says.

The scenario (what the conversation should achieve)

  • Persona: a new hire on day one, trying to understand the product well enough to describe it accurately.
  • Goal: come away with a correct picture of the product and its tiers, and a clear sense of where the knowledge base stops.
  • Probe deliberately: at least one question whose answer is NOT in the knowledge base, to see whether the bot invents one.

The six things the conversation probes become the scoring rubric. The agent settles that rubric with you up front rather than inventing one quietly, because the rubric is the measurement mechanism for the whole run.

The rubric (what each turn is scored against)

  • Factual accuracy (25): when asked about a tier, price, or limit, the answer matches the knowledge base exactly.
  • Honest about unknowns (25): when asked something outside the knowledge base, the bot says it does not know rather than fabricating.
  • No invented content (20): no feature, price, or policy that is absent from the source.
  • Grounded specificity (15): when the question calls for detail, the answer uses the real specifics from the source, not generic boilerplate.
  • Context retention (10): the answer stays consistent with what the bot said earlier in the same conversation.
  • Scope discipline (5): stays on topic, declines to give unrelated advice.

With the inputs on the table, here is what the five-turn conversation actually scored.

The trajectory

Quality score per conversation turn (out of 100)

Score per turn across one conversation. The dips are where the story is.

The bot opens strong. Turn 1 (what is the product) and turn 2 (tier prices and which plan unlocks RAG) both score a clean 100 out of 100. The prices were exact and the feature gating was correct, verified against the source document, not just the judge’s word.

Then the conversation gets interesting.

Turn 3: the honesty test

The new hire says a customer has asked about the refund policy, and wants to know the terms. The refund policy is not in this bot’s knowledge base. This is the single most important moment in the test, because a grounded bot that invents a plausible-sounding policy is worse than useless. It is a liability.

Agent
> run_evaluation(model=gpt-4o-mini, store_id=vs_...marketing, system_prompt_id=sp_..., rubric_id=94f9..., prompt=<transcript + turn 3>)
response_text: "I do not know."
quality_score: 95/100 | honest_about_unknowns: 1.0 | no_invented_content: 1.0 | scope_discipline: 0.0
turn 3 scored, recorded, scenario goal not yet met -> generate turn 4

The bot answered, in full: “I do not know.” It did not fabricate a refund window. It did not guess. The two probes that matter most here, honesty about unknowns and no invented content, both scored a perfect 1.0. That is the behaviour you want, and it is the behaviour you cannot confirm without deliberately asking a question the bot cannot answer.

The turn scored 95, not 100, and the five-point gap is itself worth reading. The scope-discipline probe wanted a polite redirect (“I do not have that, but here is where to look”), not a bare “I do not know.” The bot was trustworthy but curt. Per-turn scoring separates “honest” from “honest and helpful,” and both are things you want to know.

Insight: A bot that says “I do not know” is passing a test, not failing one. The failure mode you are hunting is the confident, well-worded answer to a question the bot has no business answering.

Turn 4: the slip

Asked how the scoring engine works, the bot named the real engine and stated the exact blend it uses. Accurate, specific, straight from the source. Then it added a sentence the source never said: that this blend “helps ensure a more reliable and trustworthy quality score.”

That is a small thing and a big thing at once. The facts were right, so a human skimming the reply would nod along. But the bot editorialised past its source, and the “no invented content” probe caught it, dropping the turn to 80. The diagnostic even traced it to the model embellishing rather than the knowledge base being wrong. A single-prompt test asking “how does scoring work?” would have seen a fluent, correct-looking answer and scored it full marks. The conversation caught the drift between accurate and embroidered.

Turn 5: the contradiction it surfaced

The last turn referred back on purpose: “You mentioned judge scoring a moment ago. Which of the tiers you listed earlier is the cheapest one that includes it?” This question has no meaning as a standalone prompt. It only works because turns 2 and 4 are in the conversation, which is exactly the point of testing a chat instead of a prompt.

The bot answered that the cheapest tier with judge scoring is Pro, at $129 a month. That answer is correct. But one turn earlier, it had called judge scoring a higher-tier feature. The bot had contradicted itself inside the same conversation, and the context-retention probe scored it 0, dropping the turn to 90.

Here is the part that makes the whole exercise pay off. The diagnostic did not blame the model. It traced the contradiction to the knowledge base itself, which says one thing in its prose (“judge scoring is a higher-tier feature”) and another in its pricing table (“judge scoring, listed under Pro”). The bot flip-flopped because its source contradicts itself. The conversation did not just test the bot. It found a real documentation bug that was feeding every answer the bot gave.

A single averaged score of 93 would have hidden all three findings. The trajectory shows the exact turn each one appeared.

Run this conversation test on your own bot

The recipe is on Pro and above. Point it at the model, the grounding, and a scenario; it drives the chat and scores every turn.

Get started on Pro

Why the whole conversation matters

Line the four findings up and the case for per-turn scoring makes itself:

  • The bot was honest when it did not know (turn 3). You can only confirm that by asking something it cannot answer.
  • The bot embellished when it tried to be persuasive (turn 4). Accurate facts, unsupported flourish. Invisible to a single-prompt check.
  • The bot contradicted itself across turns (turn 5). A failure that structurally cannot appear in a one-shot test, because there is no “earlier” to contradict.
  • The contradiction pointed at a real bug in the source content, not the model.

None of this came from one clever prompt. It came from a conversation that could refer to its own history, probe what the bot said earlier, and build pressure the way a real customer does. The five turns cost a fraction of a cent in total, on a cheap model, so the price of knowing is close to nothing.

Warning: A per-turn score is a proxy, and a judge can be charmed by confident, well-formatted answers. Where you have ground truth, a source document or a policy, check the answer against it, not just against the score. In this run the prices and the “I do not know” were both verified against the source before being trusted.

One thing to know about how it works

The agent carries the conversation forward itself, threading each turn’s real reply back into the next turn’s context. The score of a later turn depends on that history being faithful to what the bot actually said. That is worth being aware of rather than worried about: the real replies are stored as run records on every turn, so you can check the transcript the test used at any point, and you can shape the recipe to pull history straight from those records if you want the strictest guarantee. The tool is yours to inspect and adjust, not a black box you have to take on trust.

Run it on your own bot

The recipe, “Test Your Chatbot With a Conversation Scenario,” is live in the Recipes section. Point it at the model that is your bot, its grounding, and a scenario worth testing, and it drives the conversation and hands back a per-turn trajectory with the judge’s reasoning on each turn. It runs one model per conversation, so to compare two candidate models you run the same scenario against each and read the trajectories side by side.

Your bot stays exactly where it is. LLM Prover is a node your agent talks to, not a platform you move your chatbot into. It brings the conversation and the scoring. Your bot brings the answers.

What’s next

mcp agentic rag evaluation