MCP Recipes
Pro+Set Up a RAG Store and Verify It Works
The agent creates a vector store over MCP and smoke-tests retrieval once it's ready; you upload the documents in the dashboard (MCP can't send files). A human-in-the-loop setup that catches silent ingest failures.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, set up a RAG vector store and prove it actually retrieves.
THIS IS A HUMAN-IN-THE-LOOP RECIPE: you create and verify the store over MCP, but file
upload is NOT possible over MCP (MCP is JSON and cannot carry files) -- the human uploads
in the dashboard. Be explicit about that hand-off; do not pretend you can upload.
1. Confirm the tools you need are available (list the tools).
2. Ask me what the store is for (so you can name it well) and whether it is a general
knowledge store or a compliance store (regulations). Create it: create_rag_store(name,
[description], store_type). Note the store_id it returns.
3. HAND OFF TO ME CLEARLY, as a direct request-and-wait. Say something like: "I've created
the store '<exact name>' (store_id: <store_id>). I can't upload the files myself -- MCP
can't send files -- so please upload them for me: open the dashboard, go to the Stores
view, find the store named '<exact name>', and upload <which documents>. Let me know
once it's done and I'll validate it worked (check it synced and smoke-test retrieval)."
Give me:
- the exact store NAME and its store_id, so I know precisely which store to act on,
- the dashboard location (Stores view) and which documents to upload,
- a clear statement that you are waiting on me and what you'll do next when I confirm.
Then STOP and wait for me to tell you I've finished uploading. Do NOT poll in a loop
while you wait, and do NOT proceed to verify an empty store.
4. When I tell you I'm done, VERIFY the store is actually usable -- this is the point of the
recipe: call get_rag_store(store_id) and check its sync_status. If "syncing", tell me to
wait and check again shortly. If "dirty", STOP -- ingestion failed or paused, the store
is not usable; tell me to fix it, do not proceed to a paid smoke-test against it. Only
treat the store as ready when sync_status is "ready". Also sanity-check chunk_count
against file_count: a file that produced very few chunks (e.g. 1-2 chunks for a
multi-page document) is a silent ingest problem -- flag it, because retrieval will be
poor even though the store looks "ready". (get_rag_store returns sync_status, file_count,
and chunk_count for the one store.)
5. SMOKE-TEST retrieval -- do not declare success on status alone. ONLY run this if step 4
confirmed sync_status is "ready" -- never smoke-test (a paid call) against a syncing or
dirty store, you would just burn spend retrieving broken context. Ask me for a sample
question whose answer is in the uploaded documents, then run_comparison(prompt=<that
question>, models=<one model from list_models>, store_id=<the store_id>). Poll to
completion. Confirm the answer actually reflects the document content (retrieval worked),
not general knowledge. If the answer ignores the documents, the store is not retrieving
usefully -- tell me, and suspect the chunking (see step 4).
6. Report the outcome: store name + store_id, sync_status, file/chunk counts, and whether
the smoke-test proved retrieval. Tell me the store_id is what I pass as store_id (or
compliance_store_id for a compliance store) in comparisons, evaluations, and benchmarks.
ON ANY FAILURE or clear mismatch (store will not create, stays "dirty", chunk_count looks
wrong, smoke-test shows the documents were ignored): tell me honestly what happened and at
which step -- do not declare the store ready when it is not, and never claim retrieval works
if the smoke-test did not show the documents being used.
Goal
Stand up a RAG store that is actually proven to retrieve, not just created. The agent creates the store and – after you upload the documents in the dashboard – verifies it synced and smoke-tests a real question against it, catching the silent ingest failures (a document that chunked into almost nothing) that make a store look ready but retrieve badly.
When to use
Reach for this whenever you are setting up a store you will rely on – knowledge base or compliance regulations – and want confidence it works before you build evaluations or benchmarks on top of it.
Human in the loop: the one step MCP can’t do
File upload is the single capability the dashboard has that MCP does not: MCP is a JSON protocol and cannot carry file uploads. So this recipe is a hand-off. The agent creates the store and tells you its exact name and store_id; you upload the documents in the dashboard Stores view; you tell the agent when you’re done; the agent verifies and smoke-tests. The agent is explicit about this boundary rather than pretending it can upload – and it waits for you rather than racing ahead to test an empty store.
How it works
- Create over MCP, upload in the UI. The agent does the create and all verification; the upload is yours, in the dashboard, on the store the agent names for you.
- Ready is necessary, not sufficient. The agent checks
sync_status: readyAND sanity- checks chunk_count – a multi-page document that produced 1-2 chunks ingested badly and will retrieve poorly even though status says ready. - Smoke-test proves it. The agent runs a real question whose answer is in the documents and confirms the response reflects the document content. Status alone is not success.
Verification
- A store exists with the name/store_id the agent reported to you.
- The agent waited for your “uploaded” signal rather than verifying an empty store.
- sync_status is ready and chunk_count is sane for the documents uploaded.
- A smoke-test question returned an answer grounded in the documents, not general knowledge.
Where this leads
Natural next recipes once you are comfortable with this one.
Recipes that lead here