How a language model writes, one word at a time
It doesn't plan the whole sentence. It looks at the words so far, gives every possible next word a probability, picks one — and repeats.
The model's output so far
Step 0 of 9
Possible next words
probabilities add up to 100 %
Temperature
1.00
careful · lowhigh · creative
p(w) = exp(sw ÷ T) / Σw′ exp(sw′ ÷ T)
| token | logit s | s ÷ T | p |
|---|
Low T divides the logits by a small number, so the largest score wins almost everything — careful, predictable text. High T flattens the scores, giving weaker candidates a real chance — more creative, more risky.