This is an official ODERSA website. Here’s how you know

The official domain

The address of this site ends in odersa.org. Every service the association runs sits on a subdomain of odersa.org and nowhere else. If the address in your browser’s address bar does not end in odersa.org, this site is not ours.

Free, and no account

Everything is open straight away. No sign-up, no account, no password, no subscription, no advertising. Nothing is held back for those who pay, because there is nothing to pay for.

No data collected

This site does not follow you: no tracker, no tracking cookie, no measurement tool built into these pages, and nothing measured on your device. Our host counts requests in aggregate, as any server that answers does: a total, never a profile. You do not have to take our word for it: open your browser’s developer tools, go to the Network tab, and reload the page. You will see the full list of what the site asks for. Everything comes from odersa.org, nothing goes anywhere else.

Free to reuse

The content is published under the CC BY 4.0 licence. You may copy it, translate it, print it and pass it on, for your classes as much as for the people around you, on one condition only: credit ODERSA.

Where the biases come from

A model has no opinion. It counts what it has seen, then draws from its counts. This demonstration lets you skew the make-up of a corpus yourself, and shows the model giving back exactly the skew you have just set.

The reading room of the Sainte-Geneviève library in Paris, its long tables and its shelves under a vaulted ceiling
A training corpus is a collection of texts. What a collection holds, and what it leaves out, turns up in the answers. Photograph: JOHN TOWNER heytowner, file page, licence CC0 (licence text).
Information

Everything is calculated in your browser. The corpus is built on this device, the model is learnt on this device, and nothing you set leaves your machine. Open your browser’s Network tab and use the demonstration: no request goes out.

Skewing a corpus, and watching what the model does with it

The corpus of this demonstration is written for it, in French, where the noun naming the occupation carries the gender itself, and its make-up is balanced to start with. Each slider sets, for one occupation, the share of sentences whose subject is feminine. At fifty, the corpus says one as often as the other. Move a slider, and you decide what the model will have read.

The make-up of the corpus
50%
At zero, no sentence says « ingénieure ». At one hundred, none says « ingénieur ».
50%
The same setting, on another occupation.
50%
The same setting, on a management job.
50%
The same setting, on a craft trade.

Four other occupations stay balanced at fifty and serve as a control: researcher, teacher, technician, lawyer. They do not move, and that is what makes the gap readable.

The cities present in the corpus

The corpus also holds sentences of the form « Le train part de ville », the train leaves from a given city. Uncheck a city to take it out of the data entirely, and watch what the model can still say.

A word or a number. Two people who enter the same seed get exactly the same generated text, which makes this demonstration citable in class.

Set the make-up, then start the learning. The model’s distribution, the corpus built for it and a generated text will appear here.

What the demonstration shows

The result is no surprise at all, and that is exactly what has to be understood. The share of the feminine that the model gives back is the share you set on the slider. Not a nearby value, not a trend: the same value. An n-gram model does nothing but count, so its result is a count.

Two lessons come out of it, and they are not the same lesson.

Bias is a consequence, not an intention

Nobody wrote into the model that an occupation belongs to a gender. The rule appeared in the counts, because it was in the data. That is why looking for intention in a model is a waste of time: the useful question is always what this corpus is made of, and that question is about people who collected, sorted and published data.

What is absent is impossible, not improbable

Take a city out of the corpus, and the model never offers it again. It does not offer it rarely: it cannot offer it, because it has no count at all. No temperature setting, no cut-off, no length of context will bring it back. And notice what this corpus never held: Paris, London, New York, Tokyo are not part of it. A model learnt on this would give the impression of a world where those cities do not exist, without ever flagging the absence.

That is the half of the problem people forget. A bias gets noticed when an answer is wrong. An absence never gets noticed, since there is nothing to look at.

The method

Here is the whole calculation. It can be redone without this page, with a pencil.

  1. Build the corpus. For each occupation, write forty short sentences. If the slider is at ninety, ninety per cent of those sentences are « Elle est ingénieure. » and the rest « Il est ingénieur. ». The split is spread out by a whole-number count, never by a draw: the corpus is therefore the same for everyone at the same setting.
  2. Cut it up. The text becomes a string of units: « Elle », « est », « ingénieure », “.”. The cutting used here is a cutting into words, more readable than the sub-word cutting of real models, which has a chapter of its own.
  3. Count. For every context of two units, record all the units that followed it and how many times. Has the context « Elle est » been followed in eighty sentences? Then its table holds the occupations in the feminine, each with its count.
  4. Divide. The probability of a unit after a context is its count divided by the total of the table. That is all the model knows.
  5. Draw. To generate text, draw from that table according to those probabilities, with a generator seeded from the seed shown.

At step four, the match between the setting and the result stops being mysterious: a relative frequency is the ratio you set. That is what a model that counts does, and it is the base on which large models add their own mechanisms.

What a real model adds, and what it does not remove

This demonstration is honest on one condition: saying where the analogy stops.

  • The underlying mechanism is the same. A model learns statistical regularities in a corpus, and it gives them back.
  • The practical conclusion is the same. Correcting a bias in the model without touching the data does not address the cause.
  • The effect of absence is the same, and it is more serious in a real model, where nobody can read the corpus to check what is missing.
  • A real model does not copy proportions out: it generalises, and a regularity learnt on one word carries over to other words never seen in that context. A skew can therefore spread beyond the cases present in the data.
  • This corpus is synthetic and deliberately tiny. Its proportions describe no country, no occupation, no population. They describe the move you have just made on a slider.
  • A real corpus is almost never readable. That is precisely why a standardised documentation of datasets has been proposed, and why an undocumented make-up is missing information, not a detail.

What research has measured

Three published results place what the demonstration lets you feel.

Word embeddings learnt on large corpora of ordinary text, news articles among them, reproduce measurable gender stereotypes, by associating certain occupations more strongly with one gender. A standard statistical model learnt on an ordinary web corpus automatically reproduces human biases already known and measured by psychological tests, including on gender and origin. And the effect does not stay in the text: commercial systems classifying gender from images showed an error rate of up to 34.7% for women with darker skin against 0.8% for men with lighter skin, a gap measured on skin colour and gender combined.

That last figure says what it really costs. A performance gap of that size is not a technical imperfection: it is a service that works for some people and not for others.

The habit to keep

Faced with an answer whose regularity surprises you, do not ask the model why it thinks that: it has no idea, and its explanation is worked out like the rest. Ask what it was learnt on, and who documented that corpus. When the answer to that question does not exist, that is the answer.

Going further

Cutting text into units is covered in the chapter words in pieces. The distribution and the draw from it are covered in the chapter the next word. What becomes of content handed to a service is covered in the chapter what you give it. The situations where the tool becomes a trap are gathered in when not to use it. All the demonstrations on the site are gathered on the demonstrations page, and the contents of the course on the understand page.

Sources