MCP Recipes
ProCompare Models With Your Own Context
Run a prompt against several models with a block of your own context attached, so you see how each model answers grounded in your material rather than from general knowledge.
Run it: paste into your MCP agent
Highlighted parts are placeholders. Replace them with your own values before (or after) copying.
Using the LLM Prover MCP tools, compare how several models answer a question when grounded in context I supply, rather than from their own general knowledge. 1. Confirm the tools you need are available (list the tools). 2. Ask me for two things if I have not given them: the prompt (the question) and the CONTEXT text the models should answer from (a policy snippet, a doc excerpt, reference material). 3. Pick the models to compare: list_models and show me the candidates; respect my plan's model-per-comparison cap -- do not exceed it. Let me choose or pick a sensible spread. 4. Run the comparison with the context attached: run_comparison(prompt=<the question>, models=<the chosen models>, context=<my context text>). The context is injected so the models answer grounded in it. This is async -- it returns a JOB id; poll_job every 5 seconds until complete (stop after 5 minutes), then read the per-model results. 5. Summarise side by side: how each model answered, and -- importantly -- whether each actually used the supplied context or fell back to general knowledge or contradicted it. Note any model that ignored the context. 6. If my context is large or I will reuse it across many questions, mention once that a vector store (RAG) would let me attach documents and have the right pieces retrieved automatically, instead of pasting context every time. ON ANY FAILURE or clear mismatch (a model errors, the run will not complete): tell me honestly what happened and at which step -- summarise the models that did succeed and name the ones that did not.
Goal
See how different models answer your question when grounded in context you provide. You supply the prompt and a block of reference text; each model answers from it, and you get a side-by-side read on who used the context well and who didn’t.
When to use
Reach for this when the right answer depends on your material – a policy, a document excerpt, reference text – and you want to ground the models in it rather than rely on their general knowledge. It is the simplest form of context grounding: paste the text in directly.
How it works
- Context attached inline. The supplied text is injected into the request so every model answers grounded in it. Good for a one-off question against a manageable chunk of material.
- Side-by-side grounding check. The summary highlights not just the answers but whether each model actually used the context – the point of grounding is wasted if a model ignores it.
- Scales to RAG. For large documents or repeated use, a vector store retrieves the relevant pieces automatically; inline context is the starting point, not the ceiling.
Verification
- The comparison ran with the supplied context attached (not an empty/ignored context).
- The summary says, per model, whether the context was actually used.
- On partial failure, succeeding models are summarised and failures named.
Where this leads
Natural next recipes once you are comfortable with this one.