Text so far
What could come next?
step 1 of 12Bar length = how likely the model thinks each word is. The word with a tick is the one the model picks.
focused
creative
Temperature reshapes the odds: low = play it safe, high = wild and creative.
For experts
p(word) = e^(s ÷ T) ÷ Σ e^(s ÷ T)
Each candidate carries a raw score (a logit). Dividing the logits by the temperature T and running softmax turns them into probabilities that add up to 100 %. Small T sharpens the peak, large T flattens it.
We also cut the list to the five best guesses (that is top-k, k = 5) before sampling, so nonsense words never get a chance.