Interactive explainer
How a language model writes,
one word at a time
A language model reads the text so far, gives every possible next word a probability, picks one — and repeats. Step through a real example below.
The sentence, so far
complete
1.0
balanced — the model's natural confidence
Under the hood
The model outputs raw numbers called logits — the "score" column above. Softmax turns them into probabilities that add up to one:
p(word) = escore ÷ T ÷ Σ escore′ ÷ T
- Low temperature divides the scores by a small number, so the biggest score wins even more → focused, predictable text.
- High temperature flattens the scores → weaker candidates get a real chance → creative, sometimes odd text.
- Real systems often use top-k sampling: keep only the k best candidates before softmax, so the unlikely tail can never win.