The glossary
Every word this course uses, defined in plain terms, with an anchor so that a chapter can point straight to it.
The everyday vocabulary of this field is misleading, and knowing why is part of the course. Words borrowed from the human mind, “understand”, “know”, “learn”, “hallucinate”, describe a calculation here, never an inner state. A language model understands nothing, knows nothing and learns nothing in the sense a person does: it adjusts numerical parameters on a corpus, then works out a probable continuation from a text received as input. Saying that it “hallucinates” describes no disorder: it names a plausible, false result, produced by the very same calculation that produces a correct answer. This glossary keeps these words because the course and the research use them, but every definition brings them back to the calculation they really point to.
The words, in alphabetical order
Every entry carries an anchor: a chapter of the course can point straight to it. A linked word inside a definition has an entry of its own, higher or lower in this list.
- AI literacy
- The knowledge, skills and attitudes that make it possible to use an artificial intelligence system, to question it and, for some audiences, to design one, without surrendering to it blindly. A common framework published in 2026 by the OECD and the European Commission describes it in four competence areas, designed for primary and secondary education. The detail of that framework, and of the international frameworks that complete it, is gathered on the official frameworks page.
- Artefact
- A trace left by the making of an image or a sound, invisible to the eye or the ear but measurable by a calculation. The operations that enlarge an image while it is being generated leave, for instance, a regular fingerprint that can be spotted in the fine detail of the file’s frequencies. An artefact helps spot a fabricated image, but recent models leave fewer and fewer of them.
- Bias
- A systematic skew that the model reproduces, learnt in the regularities of the corpus it was trained on, never an opinion it would have formed. If the data associate an occupation with a gender more often, the model will give back the same association, to the same degree. Representation bias is one particular form of it: the one that comes from what is missing in the data rather than from what is over-present in it.
- Black box
- A system whose answers can be observed but whose internal reasoning cannot be read, even by the person who designed it. A language model is a black box in that sense: its parameters, billions of numbers adjusted during training, have no individual meaning that a person could read one by one. That does not prevent anyone from testing what the system does in practice, only from explaining why it does it in exactly that way.
- Confidence
- The assured tone of an answer, whether it is right or not. A model does not lower its voice when it makes something up: nothing in the way it puts things distinguishes a checked fact from a hallucination. The assurance of an answer is therefore never proof that it is accurate.
- Context window
- The maximum amount of text, counted in tokens, that a model can take into account at once to work out its answer. Whatever was written before that window no longer enters the current calculation, exactly as if it had never existed. A document longer than the window has to be cut or summarised before being handed to the model whole.
- Corpus
- The body of texts, images or sounds a model learns from, its only source of regularities. A dataset is a corpus organised and documented for a precise use; the word corpus refers to the raw content, before that work. What is not in the corpus does not exist for the model in any way: neither as something rare, nor as something false.
- Dataset
- A corpus organised and documented for a precise use: its composition, its origin and the way it was collected are supposed to be known, not just its content. A standardised form of documentation has been proposed for this since 2018, so that an observed bias can be traced back to its origin rather than discovered once the model is deployed. An undocumented dataset is, in practice, nothing more than a corpus.
- Deep learning
- A way of doing machine learning with a neural network made of many stacked layers. Each layer transforms the information received from the previous one a little further, which makes it possible to pick up regularities that a hand-written rule would not capture. It is the technique behind today’s language models and image generators.
- Deepfake
- An image, a video or a sound fabricated or altered by a model to make people believe in an event that never took place. Voice cloning is one form of it, applied to the voice. Human observers can already no longer reliably tell a recent synthetic face from a real one, and even judge it on average more trustworthy.
- Detector
- A tool that claims to say whether a text, an image or a voice was produced by a model. An independent evaluation of fourteen text detectors, including commercial tools sold to education, found none of them both reliable and accurate. A detector never returns a verdict: at best, a clue to be cross-checked.
- Diffusion model
- A technique that builds an image by starting from random visual noise, then removing it step by step until a plausible image appears, guided by a description in text. It is the method most used today behind image generators. Like any generated image, the result may carry an artefact that reveals how it was made, even when nothing is visible to the naked eye.
- Embedding
- The representation of a word or a token as a list of numbers, built during training so that two words used in neighbouring contexts receive neighbouring lists. Embeddings learnt on large corpora of ordinary text reproduce measurable gender stereotypes, for instance by associating certain occupations more strongly with one gender than the other. An embedding therefore says as much about the corpus it came from as about the words themselves.
- False positive and false negative
- The two ways a detector gets it wrong: the false positive accuses human content of having been produced by a model, the false negative lets content produced by a model through by calling it human. Several widely used text detectors wrongly classify the majority of texts by non-native speakers of English as machine-generated, while correctly recognising those of native speakers: that false positive does not fall at random. A tool that sets out to reduce one of those two risks almost always increases the other.
- Fine-tuning
- A second, shorter round of training on a narrower corpus, which adjusts an already trained model for a precise task or style. It starts from the parameters obtained in the first round of training rather than from scratch. A fine-tuned model remains a language model: its calculating machinery does not change, only its parameters move a little.
- Hallucination
- An answer that is made up but stated with the same confidence as a correct one. It comes out of the same calculation as any other answer: nothing in the way the model works separates a checked fact from a plausible extrapolation, and current training and evaluation procedures reward a plausible answer over an admission of uncertainty. One theoretical paper argues that a rate of hallucination would be mathematically unavoidable for this kind of model; that conclusion remains debated, other work disputing that its assumptions apply to real uses.
- Image generator
- A system that produces a new image from a description in text. The method most used for this today is the diffusion model. Like a language model, an image generator has seen no real scene that it would be copying: it assembles a plausible image from the visual regularities of its training corpus.
- Inference
- The calculation that produces an answer from an already trained model, as opposed to the training that fixed its parameters beforehand. It is what happens every time a prompt receives an answer: no parameter changes, only a calculation runs on values already fixed.
- Language model
- A system trained on text that works out, from what comes before, the most probable token to continue with. It repeats that calculation token after token to assemble a whole answer. It consults no database of facts alongside that calculation: everything it gives back comes out of the same operation, whether it is a checked fact or a hallucination.
- Machine learning
- The family of methods where a program sets its own behaviour from examples, rather than following rules written out by hand one by one. A language model and an image generator are two applications of it. Deep learning is the branch most used today.
- Memorisation
- The fact that a model can give back a fragment of its training data word for word, rather than rephrasing it. This grows with the size of the model, with the number of times the same content is duplicated in the data, and with the length of the text given as context before asking for a continuation. A model has no file where it could look that text up: the content altered its parameters, and it is that alteration which can, in some cases, bring it back almost intact.
- Metadata
- Information attached to a file and about the file itself, its origin, the date it was created, the changes it has been through, rather than about what it shows. Provenance data is one kind of metadata. Metadata can disappear with a simple screenshot or a save in another format, which makes it fragile as evidence.
- Neural network
- A calculating structure made of layers of simple units, each combining what it receives with parameters of its own. The name comes from a distant inspiration in the biological neuron, but a unit in a network does not think, any more than an isolated neuron thinks: it is the tuning of the whole, by training, that produces a useful calculation. The transformer is the architecture of this kind most used for text today.
- Open source
- A model whose code, whose parameters or both are published, so that others can inspect, reuse or check them, rather than reaching them only through a closed service. Openness says nothing on its own about the training data used: a model can publish its parameters without ever documenting its corpus. “Open source” and “documented dataset” therefore remain two distinct guarantees, rarely found together.
- Parameter
- A number inside the model, adjusted during training, that bears on its calculation. It is also called a weight. A model has billions of them, and not one of them, taken on its own, has a meaning a person could read: it is their whole that produces a useful calculation, never a single parameter.
- Probability distribution
- The list of possible tokens for continuing a text, each with a number between zero and one giving its probability, all of them adding up to one. The model works out that list at every step, then sampling picks a token out of it. Temperature, top-k and top-p are three ways of reshaping that list before drawing from it.
- Prompt
- The text given to a model as input to obtain an answer. How it is worded changes what the model gives back, since it works out its continuation from that precise text and no other. A prompt remains ordinary text: it commands nothing in the sense a program does, it steers a probable calculation.
- Provenance
- The checkable trace of a file’s origin and of its history of changes, attached to the file rather than hidden inside its content. A public technical standard, C2PA, defines how to attach that information in a signed and verifiable way. The standard itself acknowledges that this information can come away from the file, and therefore be lost with a simple screenshot: provenance protects the original file, not its copies.
- Representation bias
- The bias that comes from a group, a situation or a language being absent or under-represented in the training data, rather than from a distorted association. An image classification system showed an error rate of up to 34.7% for women with darker skin against 0.8% for men with lighter skin, a gap tied to the make-up of the learning data. What the data do not contain, the model cannot give back: it does not get it wrong, it knows nothing of it.
- Sampling
- The step that picks a token out of the probability distribution worked out by the model, rather than always taking the most probable one. Always taking the most probable one produces flat, repetitive text; drawing at random according to the probabilities, possibly reshaped by temperature, top-k or top-p, produces more natural text.
- Temperature
- A setting that softens or sharpens the probability distribution worked out by the model, before sampling. A higher temperature brings the probabilities of the different possible words closer together, which makes the choices less predictable; a lower temperature widens the gap in favour of the word that is already the most probable. It is not a physical measurement: the word is borrowed from a formula in statistical physics, with no connection to any real heat.
- Token
- The smallest unit of text a model handles, often shorter than a whole word. A rare or unknown word is then cut into several tokens made of fragments already known, which makes it possible to process it without ever having seen it as such. Each token is then converted into a number before any calculation: a model never reads letters, it reads numbers.
- Tokenisation
- The operation that cuts a text into tokens before any calculation. The method most used today comes originally from a data compression algorithm of 1994, with no connection to language when it was designed. The same content translated into different languages can call for up to fifteen times more tokens depending on the language, which makes some languages slower and more expensive to process than others.
- Top-k
- A method that restricts the draw to the k most probable tokens at each step, k being a number fixed in advance. The least probable tokens are thus set aside before sampling even happens, which avoids a wildly improbable choice without making the result entirely predictable.
- Top-p
- A method that restricts the draw to the smallest set of tokens whose cumulative probabilities reach p, rather than to a fixed number of tokens as top-k does. That set therefore changes size from one step to the next: wide when several words remain close in probability, narrow when one word clearly dominates the others. This method, also called nucleus sampling, replaces the systematic choice of the most probable word, which is flatter and more repetitive.
- Training
- The calculation, carried out once before anything is made available, which adjusts a model’s parameters so that it reflects the regularities of its corpus better and better. Once training is over, those parameters are fixed: using them afterwards to produce an answer is a separate operation, inference. Fine-tuning is a second, shorter round of training on a narrower corpus.
- Training-data extraction
- Obtaining, through an ordinary question put to the model, a passage that appeared word for word in its training corpus. Researchers have extracted gigabytes of training data from models genuinely deployed in production, by a simple method. It is the most direct evidence of memorisation: no longer a theoretical possibility, but a text found again.
- Transformer
- The neural network architecture behind today’s language models. Rather than reading a text token after token in a strict order, it compares each token with all the other tokens in the context window in one go, and weighs their importance against one another. It is that mechanism of simultaneous comparison which made it possible to train models on far larger corpora than before.
- Voice cloning
- A technique that reproduces the voice of a specific person from an audio sample, to make them say a sentence they never uttered. A United States public authority documented and penalised a real campaign using a cloned voice imitating a political figure to reach voters ahead of an election. It is a form of deepfake applied specifically to the voice.
- Watermark
- A signal deliberately placed in generated content, invisible in ordinary use, that a dedicated tool can find again. In a text, it can take the form of a slight statistical favouring of certain words at each step, a favouring too faint to see when reading but detectable by a calculation, a method already published and peer reviewed. A watermark assumes that the tool which placed it, or another that knows the method, still exists to look for it.
Seeing the ideas at work
This glossary settles the vocabulary. The course understand uses it chapter after chapter, and the demonstrations set it in motion: there you set a temperature, you skew a corpus, you watch a watermark appear.
Sources
The definitions that commit to a precise fact tie it here, term by term, without weighing each of them down.
- Detecting and Simulating Artifacts in GAN Fake Images, Xu Zhang, Svebor Karaman, Shih-Fu Chang, 2019. Artefact. The upsampling component of generated images leaves a characteristic and reproducible spectral fingerprint.
- Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification, Joy Buolamwini, Timnit Gebru, 2018. Representation bias. Commercial systems classifying gender from images showed an error rate of up to 34.7% for women with darker skin against 0.8% for men with lighter skin.
- FCC Makes AI-Generated Voices in Robocalls Illegal (FCC statement 24-17), Federal Communications Commission (FCC), 2024. Voice cloning. A United States public authority documented and penalised a real fraud campaign using a cloned voice, imitating a political figure ahead of an election.
- Testing of detection tools for AI-generated text, Debora Weber-Wulff et al., 2023. Detector. Fourteen detection tools evaluated: not one is both reliable and accurate.
- The Curious Case of Neural Text Degeneration, Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, Yejin Choi, 2020. Sampling and top-p. Systematically choosing the most probable word produces flat, repetitive text; nucleus sampling, which truncates the unreliable tail of the distribution, produces more natural text.
- Scalable Extraction of Training Data from (Production) Language Models, Milad Nasr et al., 2023. Training-data extraction. Gigabytes of training data were extracted from models genuinely deployed in production, by a simple attack.
- GPT detectors are biased against non-native English writers, Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou, 2023. False positive and false negative. Several widely used detectors wrongly classify the majority of texts by non-native speakers of English as AI-generated, while correctly identifying those of native speakers.
- A Watermark for Large Language Models, John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein, 2023. Watermark. It is possible to insert into generated text a statistical watermark invisible to the human eye, detectable afterwards by a statistical test.
- Scalable watermarking for identifying large language model outputs, Sumanth Dathathri, Abigail See, Sumedh Ghaisas, Po-Sen Huang, Rob McAdam, Johannes Welbl, et al., 2024. Watermark. A statistical watermark built into text generation was published and peer reviewed in the journal Nature.
- Why Language Models Hallucinate, Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang, 2025. Hallucination. Current training and evaluation procedures reward producing a plausible answer over admitting uncertainty.
- Hallucination is Inevitable: An Innate Limitation of Large Language Models, Ziwei Xu, Sanjay Jain, Mohan Kankanhalli, 2024. Hallucination. Debated: this work argues that a rate of hallucination is mathematically unavoidable for this kind of model; later work disputes that its assumptions apply to the real and finite uses of the models.
- AI-synthesized faces are indistinguishable from real faces and more trustworthy, Sophie J. Nightingale, Hany Farid, 2022. Deepfake. Human observers can no longer reliably tell a synthetic face from a real one, and even judge synthetic faces on average more trustworthy.
- Datasheets for Datasets, Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, Kate Crawford, 2018. Dataset. A standardised proposal exists for documenting the origin, the composition and the collection method of a dataset.
- Empowering Learners for the Age of AI: An AI Literacy Framework for Primary and Secondary Education (“AILit Framework”), OECD and European Commission, 2026. AI literacy. The common framework published by the OECD and the European Commission, four competence areas for primary and secondary education.
- Quantifying Memorization Across Neural Language Models, Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, Chiyuan Zhang, 2023. Memorisation. It grows with the size of the model, the duplication of an example in the data, and the length of the text given as context.
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings, Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, Adam Tauman Kalai, 2016. Embedding. Embeddings learnt on large corpora of ordinary text reproduce measurable gender stereotypes.
- Content Credentials: C2PA Technical Specification (version 2.4), Coalition for Content Provenance and Authenticity (C2PA), 2026. Provenance. A public technical standard defines how to attach verifiable, signed information about its origin to a file.
- C2PA Frequently Asked Questions, Coalition for Content Provenance and Authenticity (C2PA), 2026. Provenance. The standard itself acknowledges that its manifest can come away from the file, and therefore be lost with a simple screenshot.
- Distilling the Knowledge in a Neural Network, Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015. Temperature. The temperature parameter in the softmax formula softens or sharpens the probability distribution worked out by the model.
- A New Algorithm for Data Compression, Philip Gage, 1994. Tokenisation. The algorithm used for tokenisation today was originally designed as a general method of data compression, with no connection to language.
- Language Model Tokenizers Introduce Unfairness Between Languages, Aleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel Bibi, 2023. Tokenisation. The same content translated into different languages can require up to fifteen times more tokens depending on the language.
- Hierarchical Neural Story Generation, Angela Fan, Mike Lewis, Yann Dauphin, 2018. Top-k. Top-k sampling, which restricts the draw to the k most probable words at each step, serves to generate more varied text than an always deterministic choice.