This is an official ODERSA website. Here’s how you know

The official domain

The address of this site ends in odersa.org. Every service the association runs sits on a subdomain of odersa.org and nowhere else. If the address in your browser’s address bar does not end in odersa.org, this site is not ours.

Free, and no account

Everything is open straight away. No sign-up, no account, no password, no subscription, no advertising. Nothing is held back for those who pay, because there is nothing to pay for.

No data collected

This site does not follow you: no tracker, no tracking cookie, no measurement tool built into these pages, and nothing measured on your device. Our host counts requests in aggregate, as any server that answers does: a total, never a profile. You do not have to take our word for it: open your browser’s developer tools, go to the Network tab, and reload the page. You will see the full list of what the site asks for. Everything comes from odersa.org, nothing goes anywhere else.

Free to reuse

The content is published under the CC BY 4.0 licence. You may copy it, translate it, print it and pass it on, for your classes as much as for the people around you, on one condition only: credit ODERSA.

An ODERSA service

Knowing when to trust an artificial intelligence, and when not to use it at all

The AI Manual teaches how an artificial intelligence builds its answer, where its errors are, and when not to use it. The course is free, open to teenagers, working adults and teachers, and every demonstration runs on this device without sending anything.

  • Free, no account
  • No data collected
  • CC BY 4.0
A backlit laptop keyboard, most fingers resting on the home row, the U key pressed down
It all starts with typed words. What the machine then does with them is what this course is about. Photograph: Colin, file page, licence CC BY-SA 4.0 (licence text).
1 The text, cut into tokens 2 482 19 3305 77 640 12 Each token, turned into a number 3 ? The next token, predicted 4 The text, shown on screen
ODERSA diagram, reusable under the CC BY 4.0 licence.

In four steps: the text of the question is first cut into tokens, units smaller than a whole word. Each token becomes a number. From those numbers, the model works out the most probable token to continue the answer, then repeats the operation token after token. The text assembled from those predictions is the answer shown on screen.

The course

Understanding it, chapter after chapter

01

Words in pieces

An algorithm that merges the most frequent pairs of characters until it has a vocabulary, applied live to a sentence and compared across several languages.

Open

02

The next word

Watching, setting by setting, a model work out a probability distribution over the next word, then draw from it.

Open

03

Why it makes things up

Putting an opening from the corpus to the same model, then one absent from it, and seeing that nothing in its answer says which is which, except the count it never shows.

Open

04

Where the biases come from

A demonstration where you skew the make-up of a corpus and watch the model give the skew back to the exact figure, plus what the absence of a piece of data makes impossible.

Open

05

What you give it

The four distinct things that happen to content handed to a service, what research has measured about it being given back, and the questions to ask before uploading a file.

Open

06

A fabricated image

A demonstration that fabricates two images, reveals in their spectrum the regular trace of upsampling, then makes it disappear under noise.

Open

07

A fabricated voice

A speech sound fabricated twice, roughly then carefully, to compare what the ear takes in with what its spectrum reveals, and how far a little ordinary noise wipes it out.

Open

08

Watermarks

A green-list watermark placed on a text, a message hidden in the pixels of an image and then destroyed in front of you: two marking mechanisms, and how they differ from a declared provenance.

Open

09

When not to use it

Three tests for deciding whether a model belongs in a situation, the cases where it belongs in none, and why an AI detector proves nothing.

Open

In numbers

231

pages published, counted at every build of the site

7

hands-on demonstrations, running on this device

5

articles published on the blog

53

outside sources cited, each one checkable

The limits, stated upfront

What this course does not do

  • No call to any artificial intelligence service: the demonstrations are deterministic and run entirely on this device.
  • No home-made AI detector: detectors get it wrong, and this course teaches their clues and their limits.
  • No product recommended, no product condemned: this course explains how something works, it ranks no brand.