What this course rests on
Every official framework, every scientific source, every font and every corpus embedded in The AI Manual is listed here, with an address that opens and can be checked.
A false attribution would be worse than a missing one. Every source on this page was opened at its address before being cited here: nothing on it is copied from memory.
On this page
The official frameworks
The AI Manual gives shape to four artificial intelligence literacy frameworks, three of them international and one national, and to the documents that specify them or tie them to European law. Checked online on 13 August 2026.
- Empowering Learners for the Age of AI: An AI Literacy Framework for Primary and Secondary Education (“AILit Framework”), OECD and European Commission (DG Education / Digital Education Hub), with the support of CodeAI (formerly Code.org), 2026. The reference page of the framework: four competence areas and nineteen competencies for primary and secondary education, published on 17 and 18 June 2026.
- Empowering Learners for the Age of AI (full report, PDF), OECD Publishing (Paris), with the European Commission, 2026. The same framework, as a full report with its classroom examples for primary and secondary education.
- AI Competency Framework for Students, UNESCO, 2024. The framework for students: twelve competencies at three levels, understand, apply, create.
- AI Competency Framework for Teachers, UNESCO, 2024. The framework for teachers: fifteen competencies, the first global framework of its kind.
- Regulation (EU) 2024/1689 (AI Act): consolidated text, Article 4 “AI literacy”, European Union, EUR-Lex, 2024, text consolidated as at 27 July 2026. The artificial intelligence literacy obligation that falls on every provider and every deployer of a system, whatever its level of risk.
- Regulation (EU) 2026/1744 of 8 July 2026 (“Digital Omnibus on AI”): Article 1(5) amending Article 4, European Union, EUR-Lex, 2026. The regulation that turned this obligation of result into an obligation of means, in force since 27 July 2026.
- AI Literacy: Questions & Answers, European Commission, AI Office, 2025, updated on 27 July 2026. No measurable threshold is required: the effort expected is proportionate to the role and the risk of whoever applies it.
- AI4K12: Five Big Ideas in AI, AAAI and CSTA (Association for the Advancement of Artificial Intelligence / Computer Science Teachers Association), funded by the National Science Foundation (grant DRL-1846073), 2018. A United States framework, cited so that the European module stays a module beside the core, never the core itself.
The scientific and technical work, by theme
Every technical statement in this course carries a publication checked online on 13 August 2026, gathered here by theme rather than lined up in one unreadable list. The demonstration pages and the blog articles also cite, in their own footers, the publications specific to their subject.
Tokenisation
This work establishes how a text is cut into units smaller than a word, and why that cutting costs more in some languages than in others.
- Neural Machine Translation of Rare Words with Subword Units, Rico Sennrich et al., 2016.
- A New Algorithm for Data Compression, Philip Gage, 1994.
- Language Model Tokenizers Introduce Unfairness Between Languages, Aleksandar Petrov et al., 2023.
Prediction and sampling
This work shows how a model picks a word out of a probability distribution, and how temperature and the top-k or top-p cut-offs change that choice.
- The Curious Case of Neural Text Degeneration, Ari Holtzman et al., 2020.
- Distilling the Knowledge in a Neural Network, Geoffrey Hinton et al., 2015.
- Hierarchical Neural Story Generation, Angela Fan et al., 2018.
Hallucination
This work explains why a model makes things up with total confidence: a consequence of the way it is trained and evaluated, not an isolated breakdown.
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions, Lei Huang et al., 2023.
- Why Language Models Hallucinate, Adam Tauman Kalai et al., 2025.
- Hallucination is Inevitable: An Innate Limitation of Large Language Models, Ziwei Xu et al., 2024. Debated: this work argues that a certain rate of hallucination is mathematically unavoidable for a model used as a general problem solver; later work disputes that its formal assumptions apply to the real and finite uses of the models.
Bias
This work measures how the human biases present in training corpora turn up, measurably, in the models that come out of them.
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings, Tolga Bolukbasi et al., 2016.
- Semantics derived automatically from language corpora contain human-like biases, Aylin Caliskan et al., 2017.
- Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification, Joy Buolamwini, Timnit Gebru, 2018.
- Datasheets for Datasets, Timnit Gebru et al., 2018.
Memorisation of data
This work demonstrates that a model can give back fragments of its training data word for word, personal information included, by simple questioning.
- Extracting Training Data from Large Language Models, Nicholas Carlini et al., 2021.
- Quantifying Memorization Across Neural Language Models, Nicholas Carlini et al., 2023.
- Scalable Extraction of Training Data from (Production) Language Models, Milad Nasr et al., 2023.
Fabricated image
This work identifies the technical artefacts that give away a generated image, and notes that the human eye no longer sees them in recent synthetic faces.
- Leveraging Frequency Analysis for Deep Fake Image Recognition, Joel Frank et al., 2020.
- Detecting and Simulating Artifacts in GAN Fake Images, Xu Zhang et al., 2019.
- AI-synthesized faces are indistinguishable from real faces and more trustworthy, Sophie J. Nightingale, Hany Farid, 2022.
Fabricated voice
This work maps out the detection of synthetic voices and its limits, and a public authority documents a real fraudulent use of a cloned voice.
- WaveFake: A Data Set to Facilitate Audio Deepfake Detection, Joel Frank, Lea Schönherr, 2021.
- ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild, Xuechen Liu et al., 2021.
- FCC Makes AI-Generated Voices in Robocalls Illegal (FCC statement 24-17), Federal Communications Commission (FCC), 2024.
Watermarks and provenance
This work describes the two families of proof of origin, a manifest attached to the file and a statistical watermark inserted into the generated text, and the limits of each.
- Content Credentials: C2PA Technical Specification (version 2.4), Coalition for Content Provenance and Authenticity (C2PA), 2026.
- C2PA Frequently Asked Questions, Coalition for Content Provenance and Authenticity (C2PA), 2026. The standard itself acknowledges that a provenance manifest can come away from the file, and therefore be lost with a simple screenshot.
- A Watermark for Large Language Models, John Kirchenbauer et al., 2023.
- Scalable watermarking for identifying large language model outputs, Sumanth Dathathri et al., 2024. SynthID-Text, a statistical watermark published and peer reviewed in the journal Nature.
Detectors
This work shows that detectors of AI-generated text get it wrong, with a documented bias against people whose first language is not English, to the point where one of the companies that build them withdrew its own.
- GPT detectors are biased against non-native English writers, Weixin Liang et al., 2023.
- OpenAI scuttles AI-written text detector over “low rate of accuracy”, Devin Coldewey, TechCrunch, relaying an official addendum from OpenAI, 2023.
- Testing of detection tools for AI-generated text, Debora Weber-Wulff et al., 2023.
Material cost
This work puts figures on a fast-growing energy consumption, out of all proportion to an equivalent task handed to a specialised system.
- Energy and AI, International Energy Agency (IEA), 2025.
- Power Hungry Processing: Watts Driving the Cost of AI Deployment?, Alexandra Sasha Luccioni et al., 2024.
The fonts
The two fonts of the system are embedded as woff2, in site/assets/css/fonts/: no request to a font service, no CDN. Their licence was read again online on 14 August 2026, at the address of the file that carries it.
- Source Serif 4, SIL Open Font License, version 1.1, Adobe, 2014 to 2023, reserved font name “Source”. Embedded woff2 fonts, weights 500 and 600, italic 400: they carry the headings of the site.
- Public Sans, SIL Open Font License, version 1.1, The Public Sans Project Authors, 2015. Embedded woff2 fonts, weights 400, 600 and 700, italic 400: they carry the body text. It is the font of the United States public service design system, USWDS.
The SIL Open Font License allows the use, modification and redistribution of both fonts, on the single condition that they are not sold separately and that the licence stays with them.
The corpora of the demonstrations
One single third-party work is embedded in the demonstrations, and it is in the public domain. Two further texts are entirely of ODERSA origin: one serves to compare the same content from one language to another, the other to show how the composition of a corpus tilts a model.
A public domain text
Le Tour du monde en quatre-vingts jours, Jules Verne, 1873, in French. Public domain: Jules Verne died on 24 March 1905, and his work in French has been free of rights everywhere for a long time. The first 623 paragraphs of the novel, some 25,288 words, form the corpus that the toy model of this course trains on. The only edits, all of them reversible and declared in the corpus file, are typographic: apostrophes made typographic, dashes and quotation marks of the transcription removed, italic marks removed, chapter titles and table of contents set aside. A model that has read only this text from 1873 can know nothing of what has happened since: that is exactly what the demonstration brings out.
A text of our own
The same sentence, in the 7 languages that the ODERSA fleet plans to serve, is used to compare how many units the same meaning calls for from one language to another. That text is the one from ODERSA’s own trust banner, already published under the CC BY 4.0 licence: it is no third-party corpus, and it therefore raises no question of rights.
A corpus built in front of you
The demonstration on where bias comes from borrows no text: it builds its own, out of two sentence patterns, eight occupations and eight cities. You are the one who sets that composition, and that is the whole point of the page: a model cannot give back what it has never read. This corpus is written in French, like the novel from 1873, and the six translated pages each say so in their own language. It is of ODERSA origin: it is no third-party corpus, and it therefore raises no question of rights.
The engine that runs these demonstrations, draws from these corpora and works out probabilities is code written by ODERSA: it is no third-party library, and this site embeds none.
The diagrams
Every diagram on this site is an SVG drawn by hand by ODERSA, with no photograph and no image called from a third-party domain. Each is reusable under the CC BY 4.0 licence, like the rest of the site, and its caption carries that statement.
That is what lets the trust banner promise that no request leaves the device that opens this site: no image, just as no font and no script, is called from anywhere remote.
The photographs
Every photograph in The AI Manual illustrates a specific page, and its caption, under the image, already names its author. This table gathers each one with its full licence, the address that proves it on Wikimedia Commons, and the page that serves it.
| What it shows | Author | Licence | Proof on Wikimedia Commons | Served on |
|---|---|---|---|---|
| A backlit laptop keyboard, most fingers resting on the home row, the U key pressed down | Colin | CC BY-SA 4.0 | File page | Home |
| The reading room of the Sainte-Geneviève library in Paris, its long tables and its shelves under a vaulted ceiling | JOHN TOWNER heytowner | CC0 | File page | Where the biases come from |
| A technical room where dozens of network cables converge on patch panels | Brian Hankins (Bhankins) | Public domain | File page | What you give it |
| A photographer crouching at the edge of a football pitch, a long telephoto lens in hand | Ximeg | CC BY-SA 3.0 | File page | A fabricated image |
| Several microphones on stands around a drum kit, in a recording studio | dan paluska | CC BY 2.0 | File page | A fabricated voice |
| An empty lecture theatre, its rows of seats and writing tablets seen from the back of the room | Yinan Chen | Public Domain | File page | For teaching |
Two of these photographs, the keyboard on the home page and the photographer on the page A fabricated image, are under a share-alike licence, CC BY-SA. The text of The AI Manual is under CC BY 4.0, as the licence page states; those two photographs keep their own licence, they do not switch over to it. Reusing the page that carries them requires the photograph to be republished under that same share-alike licence, with its attribution, not under the terms of the rest of the site.
All the photographs on this page are embedded in this repository and served from the site’s own domain, never called from a third-party address, just like the fonts and the diagrams. That is what lets the trust banner promise that no request leaves the device that opens this site.
What this site owes to nobody else
- No dependency: not one line of code borrowed from a third-party library.
- No CDN: the fonts, the scripts and the images are all embedded in this repository.
- No artificial intelligence service called: the demonstrations are deterministic, they run entirely on the device that opens them.
- No measurement tool in these pages, no cookie, no account: nothing that identifies whoever opens this page.
This can be checked in a few seconds, in the browser’s Network tab, while the page loads: not one request leaves it.
How these credits are kept up
An out-of-date source, an address that has changed, an attribution wrongly placed: all of that can be reported, with no account and no form. The conditions for reusing this site are set out in the licence.