This is an official ODERSA website. Here’s how you know

The official domain

The address of this site ends in odersa.org. Every service the association runs sits on a subdomain of odersa.org and nowhere else. If the address in your browser’s address bar does not end in odersa.org, this site is not ours.

Free, and no account

Everything is open straight away. No sign-up, no account, no password, no subscription, no advertising. Nothing is held back for those who pay, because there is nothing to pay for.

No data collected

This site does not follow you: no tracker, no tracking cookie, no measurement tool built into these pages, and nothing measured on your device. Our host counts requests in aggregate, as any server that answers does: a total, never a profile. You do not have to take our word for it: open your browser’s developer tools, go to the Network tab, and reload the page. You will see the full list of what the site asks for. Everything comes from odersa.org, nothing goes anywhere else.

Free to reuse

The content is published under the CC BY 4.0 licence. You may copy it, translate it, print it and pass it on, for your classes as much as for the people around you, on one condition only: credit ODERSA.

What becomes of an uploaded file

Before uploading a file to a service that uses artificial intelligence, one precise question arises: can this content come back out, later, in front of someone else?

Uploading a file to a service that uses artificial intelligence rarely comes with a clear explanation of what happens to that file next. The question to ask is not “is the AI going to spy on me”, a phrase too vague to be checked. The question research has documented is more precise: can this content come back out, later, in front of someone else? A piece of schoolwork, a family photograph, a work file: uploading a file has become an everyday act, and the question of what becomes of it has not.

A model can retain whole passages

In 2021, Nicholas Carlini and his team published a direct demonstration. By simply questioning GPT-2, a language model that predates today’s conversational assistants, they extracted hundreds of text sequences reproduced word for word from its training data: names, telephone numbers, email addresses, chat exchanges, computer code. Some of those sequences appeared only once in the original data. Another finding of the study: the largest models memorise more than the smallest. None of that information was sought by any roundabout means: it simply came back out in answer to ordinary questions, put exactly as in any everyday use of the model.

What makes this risk grow

A second study, published by part of the same team in 2023, measures precisely what increases memorisation. Three factors, each measured separately: the size of the model, the repetition of the same example in the training data, and the length of the text given as context before asking for a continuation. The larger a model is, the more often an example has been repeated, the longer the context supplied, the higher the probability of finding a memorised fragment again. The authors describe a memorisation more widespread than had been thought, set to get worse with ever larger models, unless something is done to contain it.

The risk is not only theoretical

One objection is still possible: do these demonstrations target laboratory models, with no bearing on a service genuinely in use? A 2023 study, led by Milad Nasr, answers by targeting models already deployed: open models such as Pythia or GPT-Neo, semi-open ones such as LLaMA or Falcon, and a closed model reachable only through an online service, ChatGPT. It extracts gigabytes of training data from them. For ChatGPT, a model whose answers are normally constrained, the authors devise a technique that forces the model away from its usual behaviour: that technique brings training data back out at a rate 150 times higher than the one observed when the model answers normally. Constraining a model reduces that risk, it does not cancel it. Extracting gigabytes assumes no privileged access to the model: the study proceeds by simple questioning, in the same conditions as any everyday use of the service.

The question to ask before uploading a file

What these studies measure together shifts the original question. It is not about whether a model reads a file at the moment it receives it, which ordinary use already assumes. It is about whether that content can then be used to train or retrain a model, and so join the pool from which a fragment can, in the conditions measured above, come back out one day. Before uploading a file to a service that uses AI, two questions arise: does this service say, plainly, what uploaded data are used for? And does it offer a real choice to refuse their being used to train anything? The absence of an answer to either is, in itself, information.

What becomes of content handed to a service, and what that service really says about it, is set out in detail in the chapter of the course devoted to it.

What research has measured, and what it does not say:

  • The extraction of memorised training data, personal information included, by questioning a model genuinely deployed.
  • A memorisation that increases with the size of the model, the repetition of an example, and the length of the context supplied.
  • Real-time surveillance of what is written: these studies document a possible later re-emergence, not an instant reading passed to somebody else.

Sources

This article is published under the CC BY 4.0 licence: copy it, translate it, republish it, crediting ODERSA.

Back to the blog