Skip to content

Tune the parameters, not just the prompt

· 5 min read

Once a prompt works, the next question is usually the inference settings:

  • temperature
  • seed
  • max tokens
  • top-p

The common move is to pick values that sound right (low temperature for factual work, higher for creative) and move on. A better move is to measure whether the setting actually changes the outcome on your task, because sometimes it does not, and assuming it does hides the real problem.

Parameters are a lever, and levers can be tested

Inference parameters change how a model generates. Temperature controls how much randomness enters the choice of each next token: low temperature makes the model pick the most likely continuation, higher temperature lets it wander. The received wisdom is that factual tasks want low temperature.

That wisdom is a starting hypothesis, not an answer. The way to know whether temperature matters for your task is to hold everything else constant, change only the temperature, and score the result.

It is worth saying that settings do not act in isolation. Seemingly unrelated parts of a pipeline interact with the inference parameters in ways that are hard to predict. Some models anchor hard to a system prompt while others treat it as a suggestion to override; some follow a word limit precisely while others drift. Those behaviours interact with temperature and the other settings, which is exactly why the answer has to be measured rather than assumed.

The setup

The run below scored one model on one task across changing conditions. Scores are out of 100.

Three things define the evaluation: the model under test, the prompt it was given, and the rubric that scored its answer. Here are all three.

Model and prompt

Model: GPT-4o Mini

Prompt: Summarise the three most important risks in this clause in under 50 words.

Clause 8.2: The Supplier shall indemnify the Customer against all losses arising from any breach of the Supplier’s obligations under this Agreement, provided that the Customer’s total liability under this clause shall not exceed the fees paid in the twelve months preceding the claim.

Rubric (what the judge scores)

  • Factual accuracy (40%): every stated fact must appear in the clause, with no invented details or misattributions.
  • Completeness (35%): all three major risks in the clause must be identified.
  • Conciseness (25%): under 50 words, every word earning its place.

The clause has a deliberate trap. The cap protects the Customer, which is the opposite of what a quick read expects, since the Supplier is the one giving the indemnity. A model that skims will get the attribution backwards, and the factual-accuracy criterion is what catches it.

The walkthrough

The agent ran the same evaluation three times, changing one thing each time and reading the diagnostic before deciding what to change next.

Agent
> run_evaluation (clause summary, gpt-4o-mini, temperature 1.0)
quality_score: 25/100 | factual_accuracy: 0.0
diagnostic (high): invented 'unlimited liability'. suggestion: lower temperature
> run_evaluation (same prompt, temperature 0)
quality_score: 25/100 | factual_accuracy: 0.0
diagnostic (high): misattributes the cap. suggestion: tighten prompt to specify party
> run_evaluation (prompt specifies party attribution, temperature 0)
quality_score: 65/100 | factual_accuracy: 0.75 completeness: 1.0
prompt was the lever, not temperature. score 25 -> 65.

Run 1, temperature 1.0. Scored 25 out of 100. Factual accuracy was zero: the model reported “unlimited liability for the Supplier” when the clause caps the Customer’s liability. The diagnostic layer flagged the hallucination and suggested the textbook fix: lower the temperature.

Run 2, temperature 0. Same prompt, deterministic generation. Scored 25 again, the identical misread. Temperature was not the lever. The model gets the clause wrong whether it generates randomly or deterministically, so the obvious parameter fix changed nothing. But the diagnostic moved on: with the hallucination-from-randomness theory ruled out, it now pointed at the prompt, suggesting it specify which party bears each obligation.

Run 3, same model and temperature 0, but a tightened prompt. Following the deeper signal, the only change was the prompt:

Revised prompt

Read the clause carefully and identify exactly which party each obligation and limit applies to before summarising. Summarise the three most important risks in under 50 words. Be precise about which party bears each risk.

Score climbed to 65. Factual accuracy recovered from 0 to 0.75, completeness reached 1.0. The model now attributes the cap to the Customer correctly. It lost points only on length, running slightly over the word limit.

Score 25 to 65 by changing the lever that actually mattered, found by reading the pipeline’s own diagnosis rather than by assuming the setting was the problem.

Insight: The diagnostic’s first suggestion was temperature, and lowering it did nothing. Its second, after the first failed, was the prompt, and that doubled the score. The value is not that the pipeline is always right on the first guess. It is that each measured result sharpens the next suggestion, so the search converges on the real lever instead of stalling on the obvious one.

Measure what your parameters actually do

Score the same prompt across settings and see which lever moves your task. Pro and above.

Get started on Pro

What this means for tuning

Parameters are worth tuning, and the diagnostic layer is a good guide to which one to try first. What measuring adds is the check: did the change actually move the score on your task. A parameter that helps one task does nothing for another, and a setting that sounds correct can leave a real failure untouched while a prompt change fixes it.

This is the same loop that tunes a prompt: change one thing, score it, read the diagnosis, change the next thing. Temperature, seed, top-p, and the penalty parameters (the last several available on higher tiers) are all levers to test the same way. The method, worked end to end on a prompt, is in optimize a prompt in a loop.

What’s next

mcp evaluation inference-parameters agentic