The next word, never a certainty
The model does not know what it is going to write before writing it. At every word, it works out a probability for every possible candidate, then draws from them. This page runs that draw before your eyes, setting by setting: temperature, top-k and top-p.
Everything is calculated in your browser. The corpus is loaded with the page, the model is learnt on this device, and nothing you set leaves your machine. Open your browser’s Network tab and use the demonstration: no request goes out.
Working out a distribution, then drawing from it
A language model does not know what it is going to write before writing it. At every word, it works out a probability for every possible candidate: a whole list, never one settled answer. That list of probabilities is a distribution. What looks like a choice comes next: a draw from that distribution, the way you draw a card from a weighted deck.
That is why asking the same question twice does not necessarily give the same answer. As long as the distribution keeps several possible candidates, the draw picks one of them, not always the same. Here, the draw follows a seed that is visible and can be changed, so that the result stays reproducible: the same seed always gives the same word. An ordinary service hides its seed, which gives the impression of chance with no origin.
The model on this page learns on Around the World in Eighty Days, by Jules Verne, published in 1873, in the public domain and in its original French. For every string of one to three words already met in that text, it counts which word followed it and how many times. On this corpus, it currently knows … tokens, for a vocabulary of … different words, spread over … one-word contexts, … two-word contexts and … three-word contexts. The transcription has been edited for this demonstration: apostrophes made typographic, dashes and quotation marks of the transcription removed, italic marks removed, chapter titles and table of contents set aside, paragraphs of fewer than forty characters set aside.
This demonstration cuts the text into whole words, to stay readable. A real model cuts into units smaller than a word, tokens obtained by an entirely different calculation: see the chapter words in pieces. The mechanism of the distribution and the draw, though, stays the same in both cases.
What the demonstration shows
One step: the same formula, a result that changes before your eyes
Move the temperature slider, and the columns of the table move immediately: the logarithm does not change, its division by the temperature does, and the column after the softmax with it. Switch on a top-k or a top-p, and candidates kept until then move to discarded, in words, not only by the paleness of their bar. None of that changes the corpus or the counts: only the way of reading the same distribution changes.
Several steps: a text that repeats itself, or scatters
At a temperature close to zero, the draw almost always keeps the most probable candidate: the generated text falls quickly into a loop, the same few words coming back in the same order. At a temperature close to two, rare candidates become almost as probable as frequent ones: the text falls apart, plausible locally, sentence by sentence, but with no overall thread. Try both settings, on the same opening and the same seed, to see the two effects follow one another.
The method
What the “Generate the text” button does, one step at a time. The table in the section “One step” shows those same stages on the real candidates of the current opening: reading it one column after another is enough to redo it by hand, on five candidates and a calculator.
- Count. For the current context, record from the whole
corpus every word that followed it, and how many times. The starting
probability of a candidate is its count divided by the sum of the counts of
all the candidates of that context:
p = count ÷ total. - Move to the logarithm. Each probability becomes
logit = log(p), wherelogis the natural logarithm, the “ln” key on a scientific calculator. A logit is always negative or zero, since p lies between zero and one. - Add a bias, if there is one. An optional bias is added to
the logit before anything else:
biased logit = logit + bias. This page adds no bias; the chapter watermarks adds one, at this very stage. - Divide by the temperature. Each biased logit is divided by
the temperature T:
z = biased logit ÷ T. A small temperature widens the gap between the best candidate and the others; a large temperature wipes it out. - Move to the softmax. The z values become probabilities
that add up to one:
probability = exp(z) ÷ sum of the exp(z) of all the candidates. At zero temperature, this stage and the next are skipped: the candidate with the highest logit is taken directly, with no draw. - Cut at the top-k. The candidates are ranked from the most probable to the least probable. If a top-k of value k is active, only the first k stay in the running; the others are discarded.
- Cut at the top-p. If a top-p threshold is active, the model keeps the shortest prefix of that same list whose cumulative probabilities reach the threshold; the rest is discarded, including whatever a top-k cut-off may already have let through.
- Renormalise. The probabilities of the remaining candidates are divided by their own sum, to get back to a total of one: that is the final probability of each.
- Draw. A number taken between zero included and one excluded falls into that list of final probabilities, laid end to end: the candidate whose interval contains that number is the token kept. The context grows by one word, and step one starts again for the next word.
What a real model adds, and what it does not remove
This demonstration is honest on one condition: saying where the analogy stops.
What stays true in a large model
- It returns a probability distribution over the next unit, never a certainty.
- It draws from that distribution: the result depends on chance, here made visible and reproducible by a seed.
- Temperature, top-k and top-p act on its distribution through the same formula, at the same stages.
- It always answers: never a silence, never an “I don’t know” arising on its own initiative.
What this toy model does not do
- No learning by gradient descent: this model counts, it adjusts no weights.
- No word embeddings: every word is a bare label, not a point in a space of meaning.
- No attention: the context taken into account stops at three words, always the most recent ones.
- No meaning, no generalisation: a context never seen at the order asked for makes the model fall back, never guess.
What research has measured
Three published results place what the demonstration lets you feel.
Systematically choosing the most probable word produces flat, repetitive text; nucleus sampling, which truncates the unreliable part of the distribution (the top-p on this page), produces more natural text. The temperature parameter of the softmax formula softens or sharpens the probability distribution worked out by the model: a higher temperature makes the choices less predictable and more evenly spread between several possible words. Top-k sampling, which restricts the draw to the k most probable words at each step, is a method used to generate more varied text than an always deterministic choice.
Why a plausible text is not a true text
The model on this page has read only one novel, from 1873. Ask it to continue a sentence that resembles that novel, and it answers with a real context, at the right order. Ask it to continue a sentence that has no equivalent in that text, and it falls back on a shorter context, right down to returning the most frequent word in the whole corpus: it answers anyway, with the same apparent assurance, without ever signalling that it is guessing.
A plausible answer is an answer that resembles, in its form, what the corpus contains. Nothing in the calculation checks whether it matches a fact. A large model, learnt on billions of words rather than on a single novel, falls back less often and less visibly: that is the subject of the chapter why it makes things up.
Going further
Cutting text into units smaller than a word is covered in the chapter words in pieces. The cases where the model makes things up instead of answering are covered in the chapter why it makes things up. All the demonstrations on the site are gathered on the demonstrations page, and the contents of the course on the understand page.
Sources
- The Curious Case of Neural Text Degeneration Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, Yejin Choi, 2020. Establishes that systematically choosing the most probable word produces flat, repetitive text, and that nucleus sampling (top-p), which truncates the unreliable part of the distribution, produces more natural text.
- Distilling the Knowledge in a Neural Network Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015. Establishes that the temperature parameter in the softmax formula softens or sharpens the probability distribution worked out by the model, a higher temperature making the choices less predictable and more evenly spread between several possible words.
- Hierarchical Neural Story Generation Angela Fan, Mike Lewis, Yann Dauphin, 2018. Establishes that top-k sampling, which restricts the draw to the k most probable words at each step, is a method used to generate more varied text than an always deterministic choice.
- Le Tour du monde en quatre-vingts jours Jules Verne, J. Hetzel et Compagnie, 1873, public domain, in French. Transcription: Wikisource. Edited for this demonstration: apostrophes made typographic, dashes and quotation marks of the transcription removed, italic marks removed, chapter titles and table of contents set aside, paragraphs of fewer than forty characters set aside.