This is an official ODERSA website. Here’s how you know

The official domain

The address of this site ends in odersa.org. Every service the association runs sits on a subdomain of odersa.org and nowhere else. If the address in your browser’s address bar does not end in odersa.org, this site is not ours.

Free, and no account

Everything is open straight away. No sign-up, no account, no password, no subscription, no advertising. Nothing is held back for those who pay, because there is nothing to pay for.

No data collected

This site does not follow you: no tracker, no tracking cookie, no measurement tool built into these pages, and nothing measured on your device. Our host counts requests in aggregate, as any server that answers does: a total, never a profile. You do not have to take our word for it: open your browser’s developer tools, go to the Network tab, and reload the page. You will see the full list of what the site asks for. Everything comes from odersa.org, nothing goes anywhere else.

Free to reuse

The content is published under the CC BY 4.0 licence. You may copy it, translate it, print it and pass it on, for your classes as much as for the people around you, on one condition only: credit ODERSA.

The glossary

Every word this course uses, defined in plain terms, with an anchor so that a chapter can point straight to it.

The everyday vocabulary of this field is misleading, and knowing why is part of the course. Words borrowed from the human mind, “understand”, “know”, “learn”, “hallucinate”, describe a calculation here, never an inner state. A language model understands nothing, knows nothing and learns nothing in the sense a person does: it adjusts numerical parameters on a corpus, then works out a probable continuation from a text received as input. Saying that it “hallucinates” describes no disorder: it names a plausible, false result, produced by the very same calculation that produces a correct answer. This glossary keeps these words because the course and the research use them, but every definition brings them back to the calculation they really point to.

The words, in alphabetical order

Every entry carries an anchor: a chapter of the course can point straight to it. A linked word inside a definition has an entry of its own, higher or lower in this list.

AI literacy
The knowledge, skills and attitudes that make it possible to use an artificial intelligence system, to question it and, for some audiences, to design one, without surrendering to it blindly. A common framework published in 2026 by the OECD and the European Commission describes it in four competence areas, designed for primary and secondary education. The detail of that framework, and of the international frameworks that complete it, is gathered on the official frameworks page.
Artefact
A trace left by the making of an image or a sound, invisible to the eye or the ear but measurable by a calculation. The operations that enlarge an image while it is being generated leave, for instance, a regular fingerprint that can be spotted in the fine detail of the file’s frequencies. An artefact helps spot a fabricated image, but recent models leave fewer and fewer of them.
Bias
A systematic skew that the model reproduces, learnt in the regularities of the corpus it was trained on, never an opinion it would have formed. If the data associate an occupation with a gender more often, the model will give back the same association, to the same degree. Representation bias is one particular form of it: the one that comes from what is missing in the data rather than from what is over-present in it.
Black box
A system whose answers can be observed but whose internal reasoning cannot be read, even by the person who designed it. A language model is a black box in that sense: its parameters, billions of numbers adjusted during training, have no individual meaning that a person could read one by one. That does not prevent anyone from testing what the system does in practice, only from explaining why it does it in exactly that way.
Confidence
The assured tone of an answer, whether it is right or not. A model does not lower its voice when it makes something up: nothing in the way it puts things distinguishes a checked fact from a hallucination. The assurance of an answer is therefore never proof that it is accurate.
Context window
The maximum amount of text, counted in tokens, that a model can take into account at once to work out its answer. Whatever was written before that window no longer enters the current calculation, exactly as if it had never existed. A document longer than the window has to be cut or summarised before being handed to the model whole.
Corpus
The body of texts, images or sounds a model learns from, its only source of regularities. A dataset is a corpus organised and documented for a precise use; the word corpus refers to the raw content, before that work. What is not in the corpus does not exist for the model in any way: neither as something rare, nor as something false.
Dataset
A corpus organised and documented for a precise use: its composition, its origin and the way it was collected are supposed to be known, not just its content. A standardised form of documentation has been proposed for this since 2018, so that an observed bias can be traced back to its origin rather than discovered once the model is deployed. An undocumented dataset is, in practice, nothing more than a corpus.
Deep learning
A way of doing machine learning with a neural network made of many stacked layers. Each layer transforms the information received from the previous one a little further, which makes it possible to pick up regularities that a hand-written rule would not capture. It is the technique behind today’s language models and image generators.
Deepfake
An image, a video or a sound fabricated or altered by a model to make people believe in an event that never took place. Voice cloning is one form of it, applied to the voice. Human observers can already no longer reliably tell a recent synthetic face from a real one, and even judge it on average more trustworthy.
Detector
A tool that claims to say whether a text, an image or a voice was produced by a model. An independent evaluation of fourteen text detectors, including commercial tools sold to education, found none of them both reliable and accurate. A detector never returns a verdict: at best, a clue to be cross-checked.
Diffusion model
A technique that builds an image by starting from random visual noise, then removing it step by step until a plausible image appears, guided by a description in text. It is the method most used today behind image generators. Like any generated image, the result may carry an artefact that reveals how it was made, even when nothing is visible to the naked eye.
Embedding
The representation of a word or a token as a list of numbers, built during training so that two words used in neighbouring contexts receive neighbouring lists. Embeddings learnt on large corpora of ordinary text reproduce measurable gender stereotypes, for instance by associating certain occupations more strongly with one gender than the other. An embedding therefore says as much about the corpus it came from as about the words themselves.
False positive and false negative
The two ways a detector gets it wrong: the false positive accuses human content of having been produced by a model, the false negative lets content produced by a model through by calling it human. Several widely used text detectors wrongly classify the majority of texts by non-native speakers of English as machine-generated, while correctly recognising those of native speakers: that false positive does not fall at random. A tool that sets out to reduce one of those two risks almost always increases the other.
Fine-tuning
A second, shorter round of training on a narrower corpus, which adjusts an already trained model for a precise task or style. It starts from the parameters obtained in the first round of training rather than from scratch. A fine-tuned model remains a language model: its calculating machinery does not change, only its parameters move a little.
Hallucination
An answer that is made up but stated with the same confidence as a correct one. It comes out of the same calculation as any other answer: nothing in the way the model works separates a checked fact from a plausible extrapolation, and current training and evaluation procedures reward a plausible answer over an admission of uncertainty. One theoretical paper argues that a rate of hallucination would be mathematically unavoidable for this kind of model; that conclusion remains debated, other work disputing that its assumptions apply to real uses.
Image generator
A system that produces a new image from a description in text. The method most used for this today is the diffusion model. Like a language model, an image generator has seen no real scene that it would be copying: it assembles a plausible image from the visual regularities of its training corpus.
Inference
The calculation that produces an answer from an already trained model, as opposed to the training that fixed its parameters beforehand. It is what happens every time a prompt receives an answer: no parameter changes, only a calculation runs on values already fixed.
Language model
A system trained on text that works out, from what comes before, the most probable token to continue with. It repeats that calculation token after token to assemble a whole answer. It consults no database of facts alongside that calculation: everything it gives back comes out of the same operation, whether it is a checked fact or a hallucination.
Machine learning
The family of methods where a program sets its own behaviour from examples, rather than following rules written out by hand one by one. A language model and an image generator are two applications of it. Deep learning is the branch most used today.
Memorisation
The fact that a model can give back a fragment of its training data word for word, rather than rephrasing it. This grows with the size of the model, with the number of times the same content is duplicated in the data, and with the length of the text given as context before asking for a continuation. A model has no file where it could look that text up: the content altered its parameters, and it is that alteration which can, in some cases, bring it back almost intact.
Metadata
Information attached to a file and about the file itself, its origin, the date it was created, the changes it has been through, rather than about what it shows. Provenance data is one kind of metadata. Metadata can disappear with a simple screenshot or a save in another format, which makes it fragile as evidence.
Neural network
A calculating structure made of layers of simple units, each combining what it receives with parameters of its own. The name comes from a distant inspiration in the biological neuron, but a unit in a network does not think, any more than an isolated neuron thinks: it is the tuning of the whole, by training, that produces a useful calculation. The transformer is the architecture of this kind most used for text today.
Open source
A model whose code, whose parameters or both are published, so that others can inspect, reuse or check them, rather than reaching them only through a closed service. Openness says nothing on its own about the training data used: a model can publish its parameters without ever documenting its corpus. “Open source” and “documented dataset” therefore remain two distinct guarantees, rarely found together.
Parameter
A number inside the model, adjusted during training, that bears on its calculation. It is also called a weight. A model has billions of them, and not one of them, taken on its own, has a meaning a person could read: it is their whole that produces a useful calculation, never a single parameter.
Probability distribution
The list of possible tokens for continuing a text, each with a number between zero and one giving its probability, all of them adding up to one. The model works out that list at every step, then sampling picks a token out of it. Temperature, top-k and top-p are three ways of reshaping that list before drawing from it.
Prompt
The text given to a model as input to obtain an answer. How it is worded changes what the model gives back, since it works out its continuation from that precise text and no other. A prompt remains ordinary text: it commands nothing in the sense a program does, it steers a probable calculation.
Provenance
The checkable trace of a file’s origin and of its history of changes, attached to the file rather than hidden inside its content. A public technical standard, C2PA, defines how to attach that information in a signed and verifiable way. The standard itself acknowledges that this information can come away from the file, and therefore be lost with a simple screenshot: provenance protects the original file, not its copies.
Representation bias
The bias that comes from a group, a situation or a language being absent or under-represented in the training data, rather than from a distorted association. An image classification system showed an error rate of up to 34.7% for women with darker skin against 0.8% for men with lighter skin, a gap tied to the make-up of the learning data. What the data do not contain, the model cannot give back: it does not get it wrong, it knows nothing of it.
Sampling
The step that picks a token out of the probability distribution worked out by the model, rather than always taking the most probable one. Always taking the most probable one produces flat, repetitive text; drawing at random according to the probabilities, possibly reshaped by temperature, top-k or top-p, produces more natural text.
Temperature
A setting that softens or sharpens the probability distribution worked out by the model, before sampling. A higher temperature brings the probabilities of the different possible words closer together, which makes the choices less predictable; a lower temperature widens the gap in favour of the word that is already the most probable. It is not a physical measurement: the word is borrowed from a formula in statistical physics, with no connection to any real heat.
Token
The smallest unit of text a model handles, often shorter than a whole word. A rare or unknown word is then cut into several tokens made of fragments already known, which makes it possible to process it without ever having seen it as such. Each token is then converted into a number before any calculation: a model never reads letters, it reads numbers.
Tokenisation
The operation that cuts a text into tokens before any calculation. The method most used today comes originally from a data compression algorithm of 1994, with no connection to language when it was designed. The same content translated into different languages can call for up to fifteen times more tokens depending on the language, which makes some languages slower and more expensive to process than others.
Top-k
A method that restricts the draw to the k most probable tokens at each step, k being a number fixed in advance. The least probable tokens are thus set aside before sampling even happens, which avoids a wildly improbable choice without making the result entirely predictable.
Top-p
A method that restricts the draw to the smallest set of tokens whose cumulative probabilities reach p, rather than to a fixed number of tokens as top-k does. That set therefore changes size from one step to the next: wide when several words remain close in probability, narrow when one word clearly dominates the others. This method, also called nucleus sampling, replaces the systematic choice of the most probable word, which is flatter and more repetitive.
Training
The calculation, carried out once before anything is made available, which adjusts a model’s parameters so that it reflects the regularities of its corpus better and better. Once training is over, those parameters are fixed: using them afterwards to produce an answer is a separate operation, inference. Fine-tuning is a second, shorter round of training on a narrower corpus.
Training-data extraction
Obtaining, through an ordinary question put to the model, a passage that appeared word for word in its training corpus. Researchers have extracted gigabytes of training data from models genuinely deployed in production, by a simple method. It is the most direct evidence of memorisation: no longer a theoretical possibility, but a text found again.
Transformer
The neural network architecture behind today’s language models. Rather than reading a text token after token in a strict order, it compares each token with all the other tokens in the context window in one go, and weighs their importance against one another. It is that mechanism of simultaneous comparison which made it possible to train models on far larger corpora than before.
Voice cloning
A technique that reproduces the voice of a specific person from an audio sample, to make them say a sentence they never uttered. A United States public authority documented and penalised a real campaign using a cloned voice imitating a political figure to reach voters ahead of an election. It is a form of deepfake applied specifically to the voice.
Watermark
A signal deliberately placed in generated content, invisible in ordinary use, that a dedicated tool can find again. In a text, it can take the form of a slight statistical favouring of certain words at each step, a favouring too faint to see when reading but detectable by a calculation, a method already published and peer reviewed. A watermark assumes that the tool which placed it, or another that knows the method, still exists to look for it.

Seeing the ideas at work

This glossary settles the vocabulary. The course understand uses it chapter after chapter, and the demonstrations set it in motion: there you set a temperature, you skew a corpus, you watch a watermark appear.

Sources

The definitions that commit to a precise fact tie it here, term by term, without weighing each of them down.