How to find the right LLM for your use case
Selecting a model for a contract review pipeline starts with a problem: five candidates, a 14x cost spread, and no data on which one actually performs best on the specific task – extract obligations, deadlines, and risk flags from contract clauses.
The process below ran entirely from inside an IDE using LLM Prover’s MCP server. No dashboard. No manual steps. Including a rubric misalignment the pipeline caught and corrected itself.
Start with the task, not the model
The instinct is to pick a model first. GPT-4o for serious work, something cheaper for everything else. The problem is that “serious” is not a rubric. A model that excels at creative writing may miss the classification precision a contract review pipeline requires. A model that costs 12x more may score identically on your actual task.
The right starting point is a precise definition of what good looks like: the prompt, the context, and the criteria you would use to judge the output. That definition becomes the rubric. The rubric becomes the benchmark. The benchmark gives you numbers.
Step 1 – Discover the field
The first call was list_models to see what was available and at what cost.
Five models selected to span the cost and capability range:
- Grok 4.5: frontier reasoning, mid-high cost
- Claude Sonnet 4.5: strong instruction following, mid-high cost
- GPT-4.1 Mini: fast and cheap
- DeepSeek V4 Flash: very cheap, strong on structured tasks
- Qwen 3.8 27B: fast inference via Groq, mid cost
Step 2 – Run a comparison to see the responses
Before scoring anything, a comparison run across all five models returned the raw responses, costs, and latencies.
The cost spread was immediately visible: 14x between the cheapest and most expensive model. What the comparison could not tell me was which response was actually better. For that, I needed a rubric.
Step 3 – Define what good looks like
A rubric translates “good contract review” into scoreable criteria. Four dimensions, weighted by what matters in production:
- Completeness (30%) – did it identify all obligations, deadlines, and risk flags?
- Specificity (25%) – did it quote the clause text rather than paraphrase?
- Risk Depth (30%) – did it surface non-obvious risks: clause interactions, missing provisions, ambiguous terms?
- Structure (15%) – did it follow the three-section format?
The weights reflect real trade-offs. A response that quotes the clause verbatim but misses the interaction between the penalty cap and the liability cap is less useful than one that catches the interaction but paraphrases slightly.
Step 4 – Run the benchmark
With the rubric created, the benchmark ran all five models against the same prompt and scored each response.
Quality score by model -- contract clause review
Four models scored 77.5. One scored 100. Every failing model dropped on the same dimension: Risk Depth. The diagnostic layer flagged why.
Step 5 – The pipeline diagnosed itself
This is where it gets interesting.
The diagnostic judge – a separate layer that audits the scoring pipeline itself – flagged a misalignment: the rubric’s Risk Depth criterion expected clause interactions and missing provisions, but the prompt only asked for “anything that creates liability, ambiguity, or unusual exposure.” The rubric was demanding depth the prompt had not asked for. Four models scored 0.25 on Risk Depth not because they failed, but because the brief was incomplete.
The finding came from the pipeline, not from manual review. No human had to read five responses and notice the pattern. The diagnostic layer identified the root cause and suggested the fix: update the prompt to explicitly require clause interactions and missing provisions.
Run your own model selection benchmark
Connect LLM Prover to your IDE via MCP and run the process above against your actual task. Pro and above.
Get started on ProStep 6 – Fix and rerun
The rubric was updated. The benchmark was recreated with the corrected prompt. One more run.
The scores separated. DeepSeek V4 Flash scored 100 across all four dimensions. It was the only model to correctly classify every element, identify the penalty/liability cap interaction, catch the missing force majeure provision, and flag the Services vs Software scope ambiguity. Every other model failed on Completeness.
Cost per call -- contract clause review
DeepSeek V4 Flash is also the cheapest model in the field at $0.00065 per call. Grok 4.5 costs 8x more and scored 22.5 points lower. Claude Sonnet 4.5 costs 12x more and scored identically to Grok.
Here is what the quality difference looks like on the same clause:
3. Risk flags
- One-sided delay penalty: Customer may impose penalties “at its sole discretion” after >14 calendar days – 1.5%/week up to 15% of total contract value – with no corresponding Supplier remedies stated in this clause.
- Liability cap with broad carve-out: Supplier liability capped at fees paid in the prior 12 months, but unlimited for “gross negligence or wilful misconduct.”
- Customer Default trigger: Failure to give access within five (5) business days automatically “constitute[s] a Customer Default”; consequences of Default are not defined here.
- Vague access standard: “Reasonable access” to “systems and personnel” is undefined in scope, security, timing, and personnel coverage.
3. Risk flags
- Interaction between penalty cap and liability cap: The penalty is capped at “15% of the total contract value” but the general liability cap is “total fees paid in the twelve (12) months preceding the claim.” It is unclear whether the penalty sits inside or outside the liability cap – if a delay claim is also a general damages claim, which cap prevails?
- Missing force majeure provision: No excusable delay carve-out. The Supplier is strictly liable for any delay exceeding 14 days regardless of cause.
- Asymmetric Customer Default: Failure to provide access within 5 business days “shall constitute a Customer Default” with no cure period and no stated consequences – yet the clause does not specify whether a Customer Default suspends the Supplier’s delivery obligations or excuses delay penalties.
- Ambiguity in “Services” vs “Software”: The access obligation references “the Services” but the delivery obligation references “the Software.” If the scope of Services is broader than software delivery, the access obligation may be wider than intended.
- Discretionary penalty: “At its sole discretion” means the Customer can waive or selectively enforce penalties, creating uncertainty for the Supplier about its actual financial exposure.
Grok’s response is accurate. DeepSeek’s response is more useful. It caught the penalty/liability cap interaction, the missing force majeure, the cure period gap, and a scope ambiguity nobody else flagged. Both pipelines completed without errors. One delivered more actionable output.
What this means for your pipeline
The cheapest model won. Not because the others were bad – Grok and Claude produced solid analysis – but because DeepSeek followed the classification rules more precisely on this specific task with this specific rubric.
That result is not transferable. Run the same benchmark on a summarisation task or a code review task and the ranking will change. Model selection is task-specific. The model that wins on contract review may lose on customer support triage. The only way to know is to measure against your actual task with your actual rubric.
The process above ran in a single session. The rubric is reusable. Run the benchmark again when a new model ships, when your prompt changes, or when your task distribution drifts. The pipeline will tell you if something has changed. That includes whether the change is in the model or in your own rubric.
The MCP workflow shown above is available on Pro and above. For a deeper look at how agents can use LLM Prover natively, see how the MCP server works. For a real-world example of what happens when model decisions are made once and never revisited, see the multi-agent verification gap.