Skip to content

How to find the right LLM for your use case

· 8 min read

Selecting a model for a contract review pipeline starts with a problem: five candidates, a 14x cost spread, and no data on which one actually performs best on the specific task – extract obligations, deadlines, and risk flags from contract clauses.

The process below ran entirely from inside an IDE using LLM Prover’s MCP server. No dashboard. No manual steps. Including a rubric misalignment the pipeline caught and corrected itself.

Start with the task, not the model

The instinct is to pick a model first. GPT-4o for serious work, something cheaper for everything else. The problem is that “serious” is not a rubric. A model that excels at creative writing may miss the classification precision a contract review pipeline requires. A model that costs 12x more may score identically on your actual task.

The right starting point is a precise definition of what good looks like: the prompt, the context, and the criteria you would use to judge the output. That definition becomes the rubric. The rubric becomes the benchmark. The benchmark gives you numbers.

Step 1 – Discover the field

The first call was list_models to see what was available and at what cost.

Agent
> list_models
grok: grok-4.5, grok-4.6, grok-4.7 ...
anthropic: claude-haiku-4-5, claude-sonnet-4-5, claude-sonnet-5 ...
openai: gpt-4.1-nano, gpt-4.1-mini, gpt-4.1, gpt-4o ...
together: deepseek-v4-flash, deepseek-v4-pro, kimi-k3 ...
groq: qwen-3.8-27b
selected 5 models spanning cost/capability range for comparison run

Five models selected to span the cost and capability range:

  • Grok 4.5: frontier reasoning, mid-high cost
  • Claude Sonnet 4.5: strong instruction following, mid-high cost
  • GPT-4.1 Mini: fast and cheap
  • DeepSeek V4 Flash: very cheap, strong on structured tasks
  • Qwen 3.8 27B: fast inference via Groq, mid cost

Step 2 – Run a comparison to see the responses

Before scoring anything, a comparison run across all five models returned the raw responses, costs, and latencies.

Agent
> run_comparison models: grok-4.5, claude-sonnet-4-5, gpt-4.1-mini, deepseek-v4-flash, qwen-3.8-27b
job_id: 5657867a status: pending
polling ... running ... done
> get_comparison id: df05064c
grok-4.5: $0.00463 15,279ms
claude-sonnet-4-5: $0.00763 10,100ms
gpt-4.1-mini: $0.00059 5,413ms
deepseek-v4-flash: $0.00054 17,045ms
qwen-3.8-27b: $0.00219 2,320ms
cost spread: 14x. quality unknown -- need rubric scoring

The cost spread was immediately visible: 14x between the cheapest and most expensive model. What the comparison could not tell me was which response was actually better. For that, I needed a rubric.

Step 3 – Define what good looks like

A rubric translates “good contract review” into scoreable criteria. Four dimensions, weighted by what matters in production:

Agent
> create_rubric name: 'Contract clause review'
criterion: Completeness weight: 30
criterion: Specificity weight: 25
criterion: Risk Depth weight: 30
criterion: Structure weight: 15
rubric_id: 9ef33ab8 status: created
rubric maps to production requirements -- what matters in a real contract review
  • Completeness (30%) – did it identify all obligations, deadlines, and risk flags?
  • Specificity (25%) – did it quote the clause text rather than paraphrase?
  • Risk Depth (30%) – did it surface non-obvious risks: clause interactions, missing provisions, ambiguous terms?
  • Structure (15%) – did it follow the three-section format?

The weights reflect real trade-offs. A response that quotes the clause verbatim but misses the interaction between the penalty cap and the liability cap is less useful than one that catches the interaction but paraphrases slightly.

Step 4 – Run the benchmark

With the rubric created, the benchmark ran all five models against the same prompt and scored each response.

Agent
> create_benchmark rubric: 9ef33ab8 models: 5 schedule: daily
suite_id: bf119b1c status: created
> trigger_benchmark_run suite_id: bf119b1c
job_id: e7b74ad9 status: pending
polling ... running ... done
qwen-3.8-27b: 100.0 $0.00243 2,320ms
grok-4.5: 77.5 $0.00418 12,552ms
claude-sonnet-4-5: 77.5 $0.00763 9,977ms
gpt-4.1-mini: 77.5 $0.00062 3,586ms
deepseek-v4-flash: 77.5 $0.00054 17,045ms
all 4 failing models: Risk Depth score 0.25 -- diagnostic flagged rubric/prompt misalignment

Quality score by model -- contract clause review

Four models scored 77.5. One scored 100. Every failing model dropped on the same dimension: Risk Depth. The diagnostic layer flagged why.

Step 5 – The pipeline diagnosed itself

This is where it gets interesting.

The diagnostic judge – a separate layer that audits the scoring pipeline itself – flagged a misalignment: the rubric’s Risk Depth criterion expected clause interactions and missing provisions, but the prompt only asked for “anything that creates liability, ambiguity, or unusual exposure.” The rubric was demanding depth the prompt had not asked for. Four models scored 0.25 on Risk Depth not because they failed, but because the brief was incomplete.

The finding came from the pipeline, not from manual review. No human had to read five responses and notice the pattern. The diagnostic layer identified the root cause and suggested the fix: update the prompt to explicitly require clause interactions and missing provisions.

Run your own model selection benchmark

Connect LLM Prover to your IDE via MCP and run the process above against your actual task. Pro and above.

Get started on Pro

Step 6 – Fix and rerun

The rubric was updated. The benchmark was recreated with the corrected prompt. One more run.

Agent
> update_rubric id: 9ef33ab8
Risk Depth: added explicit examples -- clause interactions, missing provisions, asymmetric obligations
rubric updated
> delete_benchmark suite_id: bf119b1c
deleted
> create_benchmark prompt: updated -- now explicitly requires clause interactions and missing provisions
suite_id: b41bc5ea
> trigger_benchmark_run suite_id: b41bc5ea
polling ... running ... done
deepseek-v4-flash: 100.0 $0.00065 53,931ms -- pipeline_health: clean
grok-4.5: 77.5 $0.00511 19,868ms -- completeness: 0.25
claude-sonnet-4-5: 77.5 $0.00773 11,062ms -- completeness: 0.25
qwen-3.8-27b: 77.5 $0.00369 2,273ms -- completeness: 0.25, truncated output
gpt-4.1-mini: 55.0 $0.00071 3,455ms -- completeness: 0.25, risk_depth: 0.25
winner: deepseek-v4-flash score: 100 cost: $0.00065 -- cheapest model, highest quality

The scores separated. DeepSeek V4 Flash scored 100 across all four dimensions. It was the only model to correctly classify every element, identify the penalty/liability cap interaction, catch the missing force majeure provision, and flag the Services vs Software scope ambiguity. Every other model failed on Completeness.

Cost per call -- contract clause review

DeepSeek V4 Flash is also the cheapest model in the field at $0.00065 per call. Grok 4.5 costs 8x more and scored 22.5 points lower. Claude Sonnet 4.5 costs 12x more and scored identically to Grok.

Here is what the quality difference looks like on the same clause:

Grok 4.5 -- score: 77.5 cost: $0.00511

3. Risk flags

  • One-sided delay penalty: Customer may impose penalties “at its sole discretion” after >14 calendar days – 1.5%/week up to 15% of total contract value – with no corresponding Supplier remedies stated in this clause.
  • Liability cap with broad carve-out: Supplier liability capped at fees paid in the prior 12 months, but unlimited for “gross negligence or wilful misconduct.”
  • Customer Default trigger: Failure to give access within five (5) business days automatically “constitute[s] a Customer Default”; consequences of Default are not defined here.
  • Vague access standard: “Reasonable access” to “systems and personnel” is undefined in scope, security, timing, and personnel coverage.
DeepSeek V4 Flash -- score: 100 cost: $0.00065

3. Risk flags

  • Interaction between penalty cap and liability cap: The penalty is capped at “15% of the total contract value” but the general liability cap is “total fees paid in the twelve (12) months preceding the claim.” It is unclear whether the penalty sits inside or outside the liability cap – if a delay claim is also a general damages claim, which cap prevails?
  • Missing force majeure provision: No excusable delay carve-out. The Supplier is strictly liable for any delay exceeding 14 days regardless of cause.
  • Asymmetric Customer Default: Failure to provide access within 5 business days “shall constitute a Customer Default” with no cure period and no stated consequences – yet the clause does not specify whether a Customer Default suspends the Supplier’s delivery obligations or excuses delay penalties.
  • Ambiguity in “Services” vs “Software”: The access obligation references “the Services” but the delivery obligation references “the Software.” If the scope of Services is broader than software delivery, the access obligation may be wider than intended.
  • Discretionary penalty: “At its sole discretion” means the Customer can waive or selectively enforce penalties, creating uncertainty for the Supplier about its actual financial exposure.

Grok’s response is accurate. DeepSeek’s response is more useful. It caught the penalty/liability cap interaction, the missing force majeure, the cure period gap, and a scope ambiguity nobody else flagged. Both pipelines completed without errors. One delivered more actionable output.

What this means for your pipeline

The cheapest model won. Not because the others were bad – Grok and Claude produced solid analysis – but because DeepSeek followed the classification rules more precisely on this specific task with this specific rubric.

That result is not transferable. Run the same benchmark on a summarisation task or a code review task and the ranking will change. Model selection is task-specific. The model that wins on contract review may lose on customer support triage. The only way to know is to measure against your actual task with your actual rubric.

The process above ran in a single session. The rubric is reusable. Run the benchmark again when a new model ships, when your prompt changes, or when your task distribution drifts. The pipeline will tell you if something has changed. That includes whether the change is in the model or in your own rubric.


The MCP workflow shown above is available on Pro and above. For a deeper look at how agents can use LLM Prover natively, see how the MCP server works. For a real-world example of what happens when model decisions are made once and never revisited, see the multi-agent verification gap.

model-selection benchmarking mcp tutorial