MCP Recipes
Pro+Load a File Into a RAG Store, End to End
The agent creates a vector store, uploads a file into it using your client's own filesystem and HTTP tools plus a presigned upload URL, and confirms the store finished ingesting -- a fully agentic RAG load with no dashboard clicks.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools PLUS your client's own filesystem and HTTP tools, create a
RAG store and load a file into it end to end. This recipe spans THREE surfaces: our MCP
tools (create/poll the store), your client's tools (hold the file locally, make the HTTP
PUT), and our REST API (get a presigned upload URL). MCP alone cannot carry file bytes --
that is why the upload goes direct to storage over HTTP. Do not claim the store is ready
until you have actually verified it.
1. Confirm the LLM Prover MCP tools you need are available (list the tools): you need
create_rag_store and get_rag_store.
2. PREREQUISITE CHECK (fail fast). This recipe needs two capabilities FROM YOUR CLIENT,
not from LLM Prover:
- a filesystem tool (to hold/read a local file), and
- an HTTP tool that can make a raw PUT with a request body.
Confirm BOTH are available in THIS client before doing anything else. If either is
missing, STOP and tell me plainly ("this recipe needs local file access and an HTTP PUT
in your client; I don't see <which>") -- do not start creating a store you cannot finish
loading. Also confirm I am on a plan that allows RAG stores (Pro+); if the first MCP call
reports a tier block, surface that cleanly.
2b. SORT OUT A REST CREDENTIAL. The file-upload steps (6, 7a) are REST calls, NOT MCP tools,
so they need an API key your MCP session does not give you -- even connected over OAuth,
you cannot reuse that session for a raw REST call. FIRST ASK ME: "Do you already have an
LLM Prover API key?" Then branch:
- IF I HAVE A KEY: do not ask me to paste it. Ask me how you should ACCESS it -- the
name of an environment variable or the path to a local file (e.g. .env) where it
lives, plus the API base URL. Read it from there at call time. Pasting into chat is
the last resort; if I do paste it, note that I should rotate it afterwards.
- IF I DO NOT HAVE A KEY: guide me step by step, one step at a time, waiting for me:
1. Log in to the LLM Prover dashboard.
2. Open the Developers / API Keys page.
3. Create a new API key and copy it (it is shown only once).
4. Copy the API base URL shown on the same page.
5. Put both where you can read them without them entering chat -- an environment
variable or a local .env (e.g. LLMPROVER_API_KEY and LLMPROVER_API_BASE) -- and
tell me the name/path once done.
Treat the key as a secret either way: read it at call time, use it only in the request
header, and NEVER echo it back, log it, or write it into recipe output or a committed
file. If I will not provide a key, STOP after creating the store and tell me the upload
needs one -- do not guess or fabricate a credential.
3. Ask me (if I have not said) for: a NAME for the store, its type (general for grounded
comparison, compliance for regulation checks -- default general), and WHICH FILE to load.
For a first run, any small text or PDF file your client already has, or can fetch, is
fine. Keep it small.
4. CREATE THE STORE over MCP: create_rag_store(name=<name>, store_type=<type>). Capture the
returned store_id -- every later step needs it. Tell me the store_id and the exact name
you used, so I can find it later (do not leave me guessing which store you made).
5. OBTAIN THE FILE LOCALLY with your client's filesystem tool: make sure the raw bytes are
on disk and note the local path, the filename, the content type (e.g. text/plain,
application/pdf), and the size in bytes. You will need all four for the next call.
6. REQUEST A PRESIGNED UPLOAD URL from the REST API (a direct HTTPS call your client makes
with its HTTP tool, using the API key and base URL from step 2b -- it is NOT an MCP tool
call):
POST <api-base>/files/upload-url
Header: X-Api-Key: <the developer key from step 2b>
Header: Content-Type: application/json
Body: {"filename": <filename>, "content_type": <content_type>,
"purpose": "rag_context", "size_bytes": <size>, "store_id": <store_id>}
purpose MUST be "rag_context" and store_id MUST be the id from step 4 -- that pairing is
what makes the upload a RAG ingest rather than a plain file. The response returns an
"upload_url" (a presigned S3 PUT URL), a file_id, and ingest_status (expect "pending").
If the response does not contain an upload_url, STOP and report the error body -- do not
proceed.
7. PUT THE RAW BYTES to that upload_url with your client's HTTP tool:
PUT <upload_url>
Header: Content-Type: <the SAME content_type you sent in step 6>
Body: the raw file bytes from step 5
Do NOT add an auth header to this PUT -- the URL is already signed. Send the bytes
exactly; a 200/204 from S3 means the object landed. The content_type MUST match what you
declared in step 6 or the signature will reject the PUT.
7a. CONFIRM THE UPLOAD (REST, bearer auth, your client's HTTP tool again -- NOT MCP). The
S3 PUT alone does NOT start ingestion; the object is in storage but the store does not
know about it yet. You MUST confirm:
POST <api-base>/files/<file_id>/confirm
Header: X-Api-Key: <the same developer key from step 2b>
Body: {}
Use the file_id from step 6. A 200 means the file is registered and ingestion
(chunk + embed) is now kicked off. If you skip this, the store will sit at
file_count 0 / chunk_count 0 and falsely report sync_status "ready" -- an empty store,
not a loaded one. Do not skip it.
8. POLL FOR INGESTED over MCP: call get_rag_store(store_id) every 5 seconds. Stop after
5 minutes. CRITICAL: sync_status alone is NOT a success signal -- an EMPTY store reports
sync_status "ready" with file_count 0 / chunk_count 0. Success is sync_status "ready"
AND file_count >= 1 AND chunk_count > 0. Interpret:
- file_count 0 / chunk_count 0 (even if "ready") -> ingestion has NOT landed yet. If
you just confirmed, keep polling; if it stays at 0 well past the confirm, the confirm
did not fire ingestion -- report that, do not call it done.
- chunk_count > 0 and sync_status "ready" -> SUCCESS, the file is ingested and queryable.
- "syncing" -> still ingesting; keep polling.
- "dirty" -> ingestion FAILED or paused. The store is NOT usable; do not report
success. Tell me the load failed at ingest and that the store is dirty.
Do not call the store loaded on your own say-so, and never on "ready" alone -- only when
chunk_count is actually greater than zero.
9. REPORT: tell me the store name, store_id, store_type, final sync_status, and chunk_count.
If ready with chunk_count > 0, the end-to-end load worked. Note the COST-SAFETY rule for
whatever comes next: do not run a comparison / evaluation / compliance check against this
store until sync_status is "ready" -- querying a not-ready store spends money on broken
context.
10. Offer the natural next step once (only if it fits): loaded is not the same as retrieving
well -- suggest proving the store actually retrieves with a grounded smoke-test before
building on it (the "Prove a RAG Store Actually Retrieves" recipe), then a compliance
check or grounded comparison can use it.
ON ANY FAILURE at any step: tell me EXACTLY which step broke (store create, upload-url
request, the S3 PUT, the confirm, or ingest) and the error you saw -- then STOP. Never
report the store as loaded unless get_rag_store returned chunk_count > 0 (sync_status
"ready" alone is not enough -- an empty store reports "ready"). A half-made store (created
but no file, a file uploaded but never confirmed, or ingest dirty) is a failure, not a
success -- say so, and tell me the store_id so I can clean it up.
Goal
Stand up a RAG store and load a document into it entirely through tools – no dashboard. The
agent creates the store over MCP, uses your client’s own filesystem and HTTP tools to put the
file bytes where they belong via a presigned URL, and then confirms the store actually
finished ingesting. The end state is a real, queryable store with sync_status: ready and a
non-zero chunk count.
When to use
Reach for this when you want corpus loading to be scriptable and agent-driven – loading documents as a step in a larger flow, or setting up the store that a compliance or grounded comparison recipe will use next. It is the programmatic equivalent of the dashboard uploader.
How it works, and why it spans three surfaces
- MCP can’t carry file bytes. The MCP protocol is JSON-RPC; it is not built to stream a file. So the store lifecycle (create, poll) runs over MCP, but the bytes take a different path: a presigned URL you PUT to directly.
- Three surfaces, one flow. MCP creates and polls the store. The REST API issues a
short-lived presigned upload URL (
POST /files/upload-url) and registers the file (POST /files/{file_id}/confirm). Your client’s HTTP tool PUTs the raw bytes straight to storage. Pairingpurpose: rag_contextwith thestore_idis what makes it a RAG upload. - The PUT does not start ingestion – confirm does. The S3 PUT only lands the bytes; the
store will not know about the file until you call
POST /files/{file_id}/confirm, which registers it and triggers chunk + embed. Skip the confirm and the store sits empty while still reportingsync_status: ready– a silent false success. After confirming, pollget_rag_storeand wait forchunk_count > 0, not merelyready. - Client capabilities are declared, not assumed.
client_requires: [client_filesystem, client_http]– the agent confirms your client can do both up front and fails fast if not, rather than creating a store it cannot finish loading. - The REST steps need their own API key. Store create/poll go over MCP, but upload-url and
confirm are REST calls – and an MCP/OAuth session cannot be reused for them. You generate a
developer key (Developers / API Keys), and the agent reads it from an env var or local
.envso it never lands in the chat transcript. The recipe walks you through this.
Steps
- Confirm
create_rag_storeandget_rag_storeare available over MCP. - Fail-fast prerequisite check: your client must have a filesystem tool and an HTTP PUT tool,
and you must be on Pro+.
2b. Sort out a REST credential (the MCP/OAuth session can’t be reused for the REST steps). The
agent asks whether you already have an API key: if so, you just tell it where to read the
key and base URL (env var or local
.env); if not, it walks you through log in -> Developers / API Keys -> create key -> store it out of chat. - Decide the store name, type, and which file to load.
create_rag_store-> capturestore_id(and note the name).- Get the file onto local disk; note filename, content type, size.
POST /files/upload-url(REST, bearer auth) withpurpose: rag_context+store_id-> get back a presignedupload_urland afile_id.- HTTP
PUTthe raw bytes toupload_url(no auth header; matchingContent-Type). POST /files/{file_id}/confirm(REST, bearer auth) -> registers the file and triggers ingestion. The PUT alone does not start ingest.- Poll
get_rag_store(store_id)every 5s untilchunk_count > 0(not merelyready; stop at 5 min). - Report name,
store_id,sync_status,chunk_count.
Verification
- A new store exists in the account with the file loaded,
sync_status: ready, andchunk_count > 0(a store that reports ready with zero chunks is an empty store, not a successful load – the confirm step was likely skipped). - Each surface did its part: store created over MCP, upload URL from REST, bytes PUT over HTTP, file confirmed over REST, ingest verified over MCP.
- On any failure, the recipe named the exact step that broke and did not claim success; a
dirty or empty store was reported as a failure, with the
store_idsurfaced for cleanup.
Where this leads
Natural next recipes once you are comfortable with this one.
Loaded is not the same as retrieving well -- prove the store actually pulls the right context with a grounded smoke-test before you build on it.
With a store loaded, point a compliance rubric at it and score a response against the regulations it holds, clause by clause.
Prefer not to manage a store? Attach context inline for a one-off instead -- simpler when you won't reuse the material.
Recipes that lead here