Interactive explainer · no real AI inside
How a language model writes, one word at a time
An AI never writes a whole sentence at once. It looks at everything so far, gives every possible next word a probability, picks one — then asks the same question again. Step through a real example below.
Tip: the ← and → arrow keys also step through the sentence.
Before any picking happens, the model gives every candidate word a raw score, called a logit (z). A bigger logit means the model likes that word more. These raw scores are what the bars start from.
To turn logits into probabilities, the scores are squashed through the softmax function, with a temperature T:
p(wordi) = ezi / T / Σj ezj / T
Dividing z by T is the whole trick: a small T stretches the differences apart, so the favourite dominates. A large T squashes them together, so nearly everything becomes equally likely. Try T = 0.3 versus T = 2.0 above and watch the bars.
Real systems also use top-k (keep only the k highest-scoring words) or top-p, and throw everything else away before sampling. With top-k = 1 the model always takes the winner — no creativity at all.
| word | logit z | z ÷ T | softmax p |
|---|