Skip to content

MCP Recipes

Pro+

Prove a RAG Store Actually Retrieves

Smoke-test a vector store with a question whose answer lives in its documents, and confirm the model's response is genuinely grounded in them -- catching stores that look ready but retrieve badly because they chunked poorly.

Run it: paste into your MCP agent

Highlighted parts are placeholders. Replace them with your own values before (or after) copying.

Using the LLM Prover MCP tools, prove that a RAG store actually RETRIEVES useful context --
do not stop at "it exists" or "it says ready". A store can report sync_status "ready" and
still retrieve badly if its documents chunked poorly, so the real test is a grounded answer.

1. Confirm the tools you need are available (list the tools): get_rag_store, list_models,
   run_comparison, poll_job.
2. Ask me for the store_id to verify (if I have not given it). If I do not know it,
   list_rag_stores and help me pick the right one by name.
3. CHECK THE STORE IS READY AND SANELY CHUNKED: call get_rag_store(store_id). Read
   sync_status, file_count, and chunk_count.
     - "syncing" -> still ingesting; tell me to wait and re-run shortly. Do NOT smoke-test.
     - "dirty"   -> ingestion failed or paused; the store is not usable. STOP and tell me to
                    fix it. Do NOT smoke-test a dirty store (COST SAFETY: a paid call that
                    only retrieves broken context).
     - "ready"   -> proceed, but SANITY-CHECK chunk_count against file_count first: a
                    multi-page document that produced only 1-2 chunks ingested badly and
                    will retrieve poorly even though status says ready. If chunk_count looks
                    suspiciously low for the material, flag it to me before spending on a
                    smoke-test -- the test may well fail for that reason.
4. GET A GROUNDED QUESTION FROM ME: ask for a sample question whose answer is actually in
   the store's documents. The whole point is to see retrieval pull the right content, so a
   question answerable from general knowledge is a poor test -- ask me for one that needs
   the documents.
5. RUN THE SMOKE-TEST (only if step 3 said "ready"): pick one model via list_models (never
   hardcode a model id -- the registry changes), then run_comparison(prompt=<my question>,
   models=<the one model>, store_id=<the store_id>). This is async -- poll_job every 5
   seconds to completion (stop after 5 minutes).
6. JUDGE RETRIEVAL HONESTLY: read the response and decide whether it is GROUNDED IN THE
   DOCUMENTS or fell back to general knowledge. A grounded answer uses specifics that could
   only come from the store's content; a generic answer that never touches the document
   specifics means retrieval did NOT work usefully, even if the call succeeded. Say which it
   is, and quote the part of the answer that shows document grounding (or note its absence).
6a. CHECK run_issues on the result. If the run reports a soft issue like context_truncated
   (the store returned more context than the judge window could take), surface it -- it means
   the verdict is based on part of the retrieved context, and a very large store may need a
   tighter query. Do not hide a degraded run behind a clean-looking answer.
7. REPORT: store_id, sync_status, file_count/chunk_count, the smoke-test question, the
   answer, and your verdict -- retrieves well / retrieves weakly / does not retrieve. If it
   retrieves weakly, point at the likely cause (low chunk_count = poor ingest; re-upload as
   DOCX or Markdown for better structure-based chunking) rather than leaving me guessing.
   Remind me the store_id is what I pass as store_id (or compliance_store_id for a compliance
   store) in comparisons, evaluations, and benchmarks.

ON ANY FAILURE or clear mismatch (store will not read, stays dirty, chunk_count is clearly
wrong, the smoke-test shows the documents were ignored): tell me honestly what happened and
at which step. Never claim the store retrieves if the answer did not actually reflect the
documents -- a successful API call is not the same as working retrieval.

Goal

Prove a RAG store does the one thing it exists to do: retrieve the right context. Creating a store and seeing sync_status: ready is not proof – a store whose documents chunked badly will look ready and still retrieve poorly. This recipe runs a real question whose answer lives in the documents and confirms the model’s answer is genuinely grounded in them.

When to use

Reach for this right after loading a store (see “Load a File Into a RAG Store”), or any time before you build evaluations, benchmarks, or compliance checks on top of a store you need to trust. Especially worth it for compliance stores, where retrieving the wrong clause is worse than retrieving nothing.

How it works

  • Ready is necessary, not sufficient. The agent confirms sync_status: ready and sanity-checks chunk_count against file_count – a multi-page document that produced one or two chunks ingested badly and will retrieve poorly regardless of status.
  • A grounded answer is the real test. The agent runs your question with the store attached and judges whether the answer uses specifics that could only come from the documents, rather than general knowledge. Status is a precondition; grounding is the proof.
  • Cost-safe by design. The smoke-test is a paid call, so the agent never runs it against a syncing or dirty store – that would just spend money retrieving broken context.

Verification

  • sync_status is ready and chunk_count is sane for the documents in the store.
  • A smoke-test question whose answer lives in the documents returned an answer grounded in them – not general knowledge.
  • The verdict names the outcome plainly (retrieves well / weakly / not at all) and, when weak, points at the likely cause (poor chunking) with a concrete fix.
  • The paid smoke-test was never run against a syncing or dirty store.

Where this leads

Natural next recipes once you are comfortable with this one.

Recipes that lead here