MCP Recipes
Pro+Prove a RAG Store Actually Retrieves
Smoke-test a vector store with a question whose answer lives in its documents, and confirm the model's response is genuinely grounded in them -- catching stores that look ready but retrieve badly because they chunked poorly.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, prove that a RAG store actually RETRIEVES useful context --
do not stop at "it exists" or "it says ready". A store can report sync_status "ready" and
still retrieve badly if its documents chunked poorly, so the real test is a grounded answer.
1. Confirm the tools you need are available (list the tools): get_rag_store, list_models,
run_comparison, poll_job.
2. Ask me for the store_id to verify (if I have not given it). If I do not know it,
list_rag_stores and help me pick the right one by name.
3. CHECK THE STORE IS READY AND SANELY CHUNKED: call get_rag_store(store_id). Read
sync_status, file_count, and chunk_count.
- "syncing" -> still ingesting; tell me to wait and re-run shortly. Do NOT smoke-test.
- "dirty" -> ingestion failed or paused; the store is not usable. STOP and tell me to
fix it. Do NOT smoke-test a dirty store (COST SAFETY: a paid call that
only retrieves broken context).
- "ready" -> proceed, but SANITY-CHECK chunk_count against file_count first: a
multi-page document that produced only 1-2 chunks ingested badly and
will retrieve poorly even though status says ready. If chunk_count looks
suspiciously low for the material, flag it to me before spending on a
smoke-test -- the test may well fail for that reason.
4. GET A GROUNDED QUESTION FROM ME: ask for a sample question whose answer is actually in
the store's documents. The whole point is to see retrieval pull the right content, so a
question answerable from general knowledge is a poor test -- ask me for one that needs
the documents.
5. RUN THE SMOKE-TEST (only if step 3 said "ready"): pick one model via list_models (never
hardcode a model id -- the registry changes), then run_comparison(prompt=<my question>,
models=<the one model>, store_id=<the store_id>). This is async -- poll_job every 5
seconds to completion (stop after 5 minutes).
6. JUDGE RETRIEVAL HONESTLY: read the response and decide whether it is GROUNDED IN THE
DOCUMENTS or fell back to general knowledge. A grounded answer uses specifics that could
only come from the store's content; a generic answer that never touches the document
specifics means retrieval did NOT work usefully, even if the call succeeded. Say which it
is, and quote the part of the answer that shows document grounding (or note its absence).
6a. CHECK run_issues on the result. If the run reports a soft issue like context_truncated
(the store returned more context than the judge window could take), surface it -- it means
the verdict is based on part of the retrieved context, and a very large store may need a
tighter query. Do not hide a degraded run behind a clean-looking answer.
7. REPORT: store_id, sync_status, file_count/chunk_count, the smoke-test question, the
answer, and your verdict -- retrieves well / retrieves weakly / does not retrieve. If it
retrieves weakly, point at the likely cause (low chunk_count = poor ingest; re-upload as
DOCX or Markdown for better structure-based chunking) rather than leaving me guessing.
Remind me the store_id is what I pass as store_id (or compliance_store_id for a compliance
store) in comparisons, evaluations, and benchmarks.
ON ANY FAILURE or clear mismatch (store will not read, stays dirty, chunk_count is clearly
wrong, the smoke-test shows the documents were ignored): tell me honestly what happened and
at which step. Never claim the store retrieves if the answer did not actually reflect the
documents -- a successful API call is not the same as working retrieval.
Goal
Prove a RAG store does the one thing it exists to do: retrieve the right context. Creating a
store and seeing sync_status: ready is not proof – a store whose documents chunked badly
will look ready and still retrieve poorly. This recipe runs a real question whose answer lives
in the documents and confirms the model’s answer is genuinely grounded in them.
When to use
Reach for this right after loading a store (see “Load a File Into a RAG Store”), or any time before you build evaluations, benchmarks, or compliance checks on top of a store you need to trust. Especially worth it for compliance stores, where retrieving the wrong clause is worse than retrieving nothing.
How it works
- Ready is necessary, not sufficient. The agent confirms
sync_status: readyand sanity-checkschunk_countagainstfile_count– a multi-page document that produced one or two chunks ingested badly and will retrieve poorly regardless of status. - A grounded answer is the real test. The agent runs your question with the store attached and judges whether the answer uses specifics that could only come from the documents, rather than general knowledge. Status is a precondition; grounding is the proof.
- Cost-safe by design. The smoke-test is a paid call, so the agent never runs it against a syncing or dirty store – that would just spend money retrieving broken context.
Verification
sync_statusisreadyandchunk_countis sane for the documents in the store.- A smoke-test question whose answer lives in the documents returned an answer grounded in them – not general knowledge.
- The verdict names the outcome plainly (retrieves well / weakly / not at all) and, when weak, points at the likely cause (poor chunking) with a concrete fix.
- The paid smoke-test was never run against a syncing or dirty store.
Where this leads
Natural next recipes once you are comfortable with this one.
Retrieval proven? If the store holds regulations, score a response against them clause by clause.
If retrieval is weak, compare models on the same question with the context pasted inline to isolate whether the problem is the store or the model.
Recipes that lead here