Your RAG store sets the quality ceiling for your pipeline
A RAG pipeline retrieves content from a store and feeds it to the model. When the answer is wrong or vague, the instinct is to blame the model or tune the prompt. Often the real limit is the store itself: if the document does not contain the answer, no model can retrieve it, and no prompt can conjure it. The store content is the ceiling on what the pipeline can do.
That ceiling is measurable. Put two versions of a document in two separate stores, run the same model with the same prompt and the same rubric against each, and the difference in score is the effect of the content, isolated from everything else.
The setup
One customer question, asked against two stores. Everything about the run is identical except which store the model retrieves from.
The customer question
Store A holds a thin, vague version of the returns policy. Store B holds a complete one. Here is the entire content of each.
Store A (thin)
Store B (complete)
The rubric scores whether the answer is specific and correct, grounded in the policy, and complete.
Rubric (what the judge scores)
- Specific and correct (50%): gives the exact return window, who pays shipping, and refund timing. Vague answers score low; invented details score zero.
- Grounded in policy (30%): every claim is supported by the retrieved content, not general knowledge.
- Complete (20%): addresses all parts of the question, not just the easy one.
The run
Same model, same prompt, same rubric. The store is the only thing that changes.
Against store A, the model scored 50. It answered honestly:
The model is not wrong here. It correctly refused to invent details the document did not contain, which is exactly what a grounded pipeline should do. But the answer is useless to the customer: it does not say whether 40 days still qualifies, who pays shipping, or when the money arrives. The specific-and-correct criterion scored zero, because there were no specifics to retrieve. The ceiling was the document.
Against store B, the same model scored 92.5:
Correct on every point: 40 days means store credit rather than a refund, the $6.95 fee applies, store credit arrives within one business day. Nothing changed but the document behind the pipeline.
Measure what your documents do to your AI
Run the same benchmark against two versions of your store and see the delta. Pro Plus.
The diagnostic found a gap in the good document too
Store B scored 92.5, not 100, and the reason is worth reading. The grounding criterion docked a quarter point, and the diagnostic said why: the policy states the $6.95 fee is deducted “from the refund”, but says nothing about the store-credit case the customer actually falls into. The model bridged the gap reasonably, but the source document has a real terminology hole.
That is the store-content lever meeting the diagnostic lever. The score told you the pipeline improved; the diagnostic told you the document still has a specific, fixable gap. Fix that line in the policy and the ceiling rises again. This is the same loop covered in how the diagnostic improvement loop works.
What this measures that nothing else does
Content teams have never had a clean answer to a simple question: did the work we did on our documents actually make the AI better? The two-store A/B gives it. Rewrite a knowledge-base article, put the old and new versions in separate stores, run the same benchmark against each, and the score delta is the measured value of the edit. Not a hunch, a number.
It is one of fourteen levers that set a stack’s cost and quality. The full framework, and how store content interacts with model choice, prompt, and the rest, is in the practical framework for your LLM stack.
What’s next
RAG Store Guide
How to create a context store, upload documents, and get reliable retrieval in comparisons, evaluations, and benchmarks.
Scoring Guide
How LLM Prover calculates scores, what affects them, and how to use the diagnostic judge to find out why.
Benchmark Guide
Save a run configuration, schedule it on a cadence, and track quality and cost drift over time.