Temperature
In short
Temperature is a setting that controls how random an LLM's output is, from focused and predictable at low values to more varied and creative at high values.
What is temperature in an LLM?
When a large language model generates text, it doesn't pick words with certainty. At each step it computes a score for every possible next token, turns those scores into probabilities, and then samples one. Temperature is a number, usually between 0 and 2 in most APIs, that reshapes those probabilities before sampling.
Technically, the model's raw scores, called logits, are divided by the temperature before the softmax function turns them into probabilities. A low temperature such as 0.2 sharpens the distribution so the most likely token wins almost every time, giving focused, consistent answers. A high temperature such as 1.2 flattens it, so less likely tokens are picked more often, producing more diverse and surprising text, but also more mistakes and rambling. At or near 0, the model almost always chooses the single most likely token, which is called greedy decoding.
Think of it like a spice dial on a recipe generator: turned low, you get the classic dish every time; turned high, you get experimental combinations, some brilliant and some inedible. In practice, developers use low temperatures for tasks with one right answer, such as extracting data, classification, or generating code, and higher temperatures for brainstorming, creative writing, or producing several alternative drafts.
Temperature is often confused with top-p, also called nucleus sampling. Temperature changes how spread out the probabilities are, while top-p limits the choice to the smallest set of tokens whose combined probability reaches a threshold such as 0.9, and it is usually best to adjust one or the other, not both. Also, a low temperature makes output more predictable, not more truthful, and even a temperature of 0 doesn't always guarantee identical results.
Key takeaways
- Temperature scales the model's token probabilities before one token is sampled.
- Low values, around 0 to 0.3, give focused and consistent output.
- High values, around 0.8 and above, give more varied, creative, and error-prone output.
- Temperature affects randomness, not accuracy or knowledge.
- Top-p is a related sampling setting; usually tune one or the other, not both.
Example
import math
def softmax_with_temperature(logits, temperature):
scaled = [x / temperature for x in logits]
total = sum(math.exp(x) for x in scaled)
return [round(math.exp(x) / total, 3) for x in scaled]
# Raw scores for three candidate next tokens: "blue", "clear", "purple"
logits = [2.0, 1.0, 0.1]
print(softmax_with_temperature(logits, 0.2)) # [0.993, 0.007, 0.0] almost always "blue"
print(softmax_with_temperature(logits, 1.0)) # [0.659, 0.242, 0.099]
print(softmax_with_temperature(logits, 2.0)) # [0.502, 0.304, 0.194] more varietyReaders ask
What temperature should I use?
Use a low temperature, around 0 to 0.3, for tasks with a single correct answer like data extraction, classification, or code. Use a moderate to high value, roughly 0.7 to 1.0, for brainstorming and creative writing, and test with your own prompts because the best value depends on the model and task.
Does temperature 0 make an LLM deterministic?
Mostly, but not always. It makes the model pick the most likely token at each step, yet tiny numerical differences from hardware and request batching can still change the output occasionally.
What is the difference between temperature and top-p?
Temperature reshapes the whole probability distribution, making it sharper or flatter. Top-p, or nucleus sampling, instead cuts off the unlikely tail and samples only from the smallest set of tokens whose probabilities add up to a threshold, such as 0.9.
See also
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- TokenAI & Machine Learning, p. 46A token is the basic unit of text that an LLM reads and generates, usually a whole word, part of a word, or a punctuation mark, mapped to a numeric ID.
- InferenceAI & Machine Learning, p. 24Inference is the stage where a trained machine learning model is used to make predictions or generate output from new data, without changing what it learned.
- HallucinationAI & Machine Learning, p. 23A hallucination is when an AI model, such as an LLM, confidently produces information that sounds plausible but is false, invented, or unsupported by sources.
- PromptAI & Machine Learning, p. 35A prompt is the input text or instructions you give an AI model, such as an LLM, to tell it what task to perform and what kind of answer you want.
- Generative AIAI & Machine Learning, p. 20Generative AI is artificial intelligence that creates new content, such as text, images, code, or audio, based on patterns learned from existing data.
Spotted a mistake or something missing on this page?Suggest an edit