Tune the parameters, not just the prompt
Once a prompt works, the next question is usually the inference settings:
- temperature
- seed
- max tokens
- top-p
The common move is to pick values that sound right (low temperature for factual work, higher for creative) and move on. A better move is to measure whether the setting actually changes the outcome on your task, because sometimes it does not, and assuming it does hides the real problem.
Parameters are a lever, and levers can be tested
Inference parameters change how a model generates. Temperature controls how much randomness enters the choice of each next token: low temperature makes the model pick the most likely continuation, higher temperature lets it wander. The received wisdom is that factual tasks want low temperature.