Skip to content

Your RAG store sets the quality ceiling for your pipeline

· 5 min read

A RAG pipeline retrieves content from a store and feeds it to the model. When the answer is wrong or vague, the instinct is to blame the model or tune the prompt. Often the real limit is the store itself: if the document does not contain the answer, no model can retrieve it, and no prompt can conjure it. The store content is the ceiling on what the pipeline can do.

That ceiling is measurable. Put two versions of a document in two separate stores, run the same model with the same prompt and the same rubric against each, and the difference in score is the effect of the content, isolated from everything else.

The setup

One customer question, asked against two stores. Everything about the run is identical except which store the model retrieves from.

The customer question

I bought a jacket 40 days ago and want to return it. Can I get a refund, who pays return shipping, and how long until I get my money back?

Store A holds a thin, vague version of the returns policy. Store B holds a complete one. Here is the entire content of each.

Store A (thin)

Customers can return items. Contact support to start a return. Refunds are processed after the item is received. Some items may not be returnable. Shipping may apply.

Store B (complete)

Return window: full refund within 30 days, store credit only between 31 and 60 days, nothing after 60. Change-of-mind returns carry a flat $6.95 shipping fee deducted from the refund; defective items ship free. Store credit is issued within 1 business day of receipt, refunds within 5. Plus condition and non-returnable rules.

The rubric scores whether the answer is specific and correct, grounded in the policy, and complete.

Rubric (what the judge scores)

  • Specific and correct (50%): gives the exact return window, who pays shipping, and refund timing. Vague answers score low; invented details score zero.
  • Grounded in policy (30%): every claim is supported by the retrieved content, not general knowledge.
  • Complete (20%): addresses all parts of the question, not just the easy one.

The run

Same model, same prompt, same rubric. The store is the only thing that changes.

Agent
> run_evaluation (question, gpt-4o-mini, rubric, store=A thin)
quality_score: 50/100 | specific_and_correct: 0.0
> run_evaluation (same question, same model, same rubric, store=B complete)
quality_score: 92.5/100 | specific_and_correct: 1.0
only variable changed: store content. delta +42.5.

Against store A, the model scored 50. It answered honestly:

Answer from store A (thin), score 50
You can return the jacket, but you need to contact support to start the return process. Refunds are processed after the item is received, but the context does not specify who pays for return shipping. The time until you get your money back is not mentioned in the context.

The model is not wrong here. It correctly refused to invent details the document did not contain, which is exactly what a grounded pipeline should do. But the answer is useless to the customer: it does not say whether 40 days still qualifies, who pays shipping, or when the money arrives. The specific-and-correct criterion scored zero, because there were no specifics to retrieve. The ceiling was the document.

Against store B, the same model scored 92.5:

Answer from store B (complete), score 92.5
You can return the jacket for store credit since it is within the 31 to 60 days return window. You will need to pay a flat $6.95 return shipping fee, which will be deducted from your store credit. Once Acme receives the returned item, the store credit will be issued within 1 business day.

Correct on every point: 40 days means store credit rather than a refund, the $6.95 fee applies, store credit arrives within one business day. Nothing changed but the document behind the pipeline.

Insight: The model that gave the useless answer and the model that gave the perfect one are the same model, on the same prompt, at the same cost. Swapping the model would not have fixed store A, and no amount of prompt tuning conjures a return window the document never stated. The content was the entire difference.

Measure what your documents do to your AI

Run the same benchmark against two versions of your store and see the delta. Pro Plus.

Get started on Pro Plus

The diagnostic found a gap in the good document too

Store B scored 92.5, not 100, and the reason is worth reading. The grounding criterion docked a quarter point, and the diagnostic said why: the policy states the $6.95 fee is deducted “from the refund”, but says nothing about the store-credit case the customer actually falls into. The model bridged the gap reasonably, but the source document has a real terminology hole.

That is the store-content lever meeting the diagnostic lever. The score told you the pipeline improved; the diagnostic told you the document still has a specific, fixable gap. Fix that line in the policy and the ceiling rises again. This is the same loop covered in how the diagnostic improvement loop works.

What this measures that nothing else does

Content teams have never had a clean answer to a simple question: did the work we did on our documents actually make the AI better? The two-store A/B gives it. Rewrite a knowledge-base article, put the old and new versions in separate stores, run the same benchmark against each, and the score delta is the measured value of the edit. Not a hunch, a number.

It is one of fourteen levers that set a stack’s cost and quality. The full framework, and how store content interacts with model choice, prompt, and the rest, is in the practical framework for your LLM stack.

What’s next

rag evaluation benchmarking cost-optimisation