This is an official ODERSA website. Here’s how you know

The official domain

The address of this site ends in odersa.org. Every service the association runs sits on a subdomain of odersa.org and nowhere else. If the address in your browser’s address bar does not end in odersa.org, this site is not ours.

Free, and no account

Everything is open straight away. No sign-up, no account, no password, no subscription, no advertising. Nothing is held back for those who pay, because there is nothing to pay for.

No data collected

This site does not follow you: no tracker, no tracking cookie, no measurement tool built into these pages, and nothing measured on your device. Our host counts requests in aggregate, as any server that answers does: a total, never a profile. You do not have to take our word for it: open your browser’s developer tools, go to the Network tab, and reload the page. You will see the full list of what the site asks for. Everything comes from odersa.org, nothing goes anywhere else.

Free to reuse

The content is published under the CC BY 4.0 licence. You may copy it, translate it, print it and pass it on, for your classes as much as for the people around you, on one condition only: credit ODERSA.

What you give it, and what it can give back

Uploading a text, an image or a file to an artificial intelligence service sets off four different things, which terms of use often gather into a single sentence. This chapter separates them, because the risk is not the same for each.

A technical room where dozens of network cables converge on patch panels
What you upload to an online service leaves through cables, towards machines that are not ours. Photograph: Brian Hankins (Bhankins), file page, licence Public domain.

The question “is artificial intelligence spying on me?” cannot be checked: it is too vague for an answer to be true or false. The question that can be checked is narrower. Can what I have just uploaded come back out one day, in front of someone else? That question has a measured answer, and it comes down to four steps that first have to be told apart.

The four things that “giving” brings together

One and the same upload sets off as many as four operations. They are not equally unavoidable, and they do not have the same consequences.

Reading, to answer now
The content is cut into units and placed in the model’s context, from which it works out its answer. This operation is unavoidable: it is exactly what the service is being asked for. It stops with the answer.
Keeping, on the service’s servers
The content and the answer are recorded, for the history of the conversation, for handling reports, for moderation. A retention period is a decision of the service, not a property of the model.
Retraining, with this content among others
The content joins the pool of data that will be used to train a later version of the model. This is where the upload stops being an exchange and becomes a lasting contribution, often without the person having meant it to be.
Giving back, later, to someone else
A model trained on a piece of content can reproduce it word for word in answer to an ordinary question asked by another person. This is not a hypothesis: it is the result measured by the work cited below.

The first two operations are a matter of service policy, which can be read and compared. The last two are a matter of how a model works, and that is what this chapter explains.

A model retains, and that can be measured

A model has no database where the texts it has read could be found again. Its training data are not stored somewhere inside it: they have altered its parameters. That is why the word “memory” misleads. Whole passages do come back out, all the same, and research knows under what conditions.

In 2021, a team led by Nicholas Carlini extracted from a language model hundreds of sequences reproduced word for word from its training data, including personally identifiable information. The method was nothing like a workaround: these were ordinary questions. Some of the sequences recovered appeared only once in the original data.

A second study, in 2023, measures what makes this phenomenon grow. Three factors, each established separately.

  1. The size of the model. The more parameters it has, the more it memorises.
  2. The duplication of an example in the data. Content present several times comes back out more easily than content present once.
  3. The length of the text given as context before asking for a continuation. The longer the opening, the more likely the content is to be given back.

One objection is still open at this stage: these demonstrations may target laboratory models, with no bearing on a service genuinely in use. A third study, the same year, answers by targeting models already deployed in production, including a consumer model whose answers are constrained, and extracts gigabytes of training data from them by simple questioning. Constraining a model reduces that risk. It does not cancel it.

This work measures content that can be given back later, to someone else. It does not document an instant reading passed to somebody while you are typing. Confusing the two loses the only useful thing here, which is knowing which risk is taken at which moment.

Reading what a service really says

A usage policy is read by looking for three precise formulations, not for a general impression of seriousness.

What the sentence covers

“To improve our services” is the phrase that most often covers retraining without naming it. It is neither untruthful nor informative. A useful text says instead what is done with the content, and by whom.

The choice, and its default

A way to refuse retraining sometimes exists, and its form matters more than its existence. A refusal that applies by default protects everyone. A refusal you have to go and find in a setting protects only the people who know it is there, which is to say almost nobody. A policy is judged on its default, not on its options.

The period, and what is left afterwards

A retention period says nothing about the model already trained. Deleting a conversation removes a record from a server; it does not remove from a model what that content has already altered inside it. A trained model is not untrained of a piece of content by the deletion of an account.

Before uploading a file

Three questions, in this order. The absence of an answer to any one of them is itself an answer.

  1. Does this service say, plainly and with no wrapping phrase, what uploaded content is used for?
  2. Is refusing retraining the default behaviour, or a setting you have to go and find?
  3. Does this content belong to somebody else? A work file, a pupil’s homework, a medical document, a family photograph all involve people who have agreed to nothing.

That last point is the one most often forgotten, and the only one that cannot be put right. Accepting terms for yourself is a choice. Accepting them for somebody else who is not there is not.

What can be uploaded, and what cannot

There is no universal list, because the risk depends on what is at stake. This split holds for a service whose retraining policy is not established.

  • A text that is already public, or meant to become so.
  • Content that concerns you alone and whose disclosure would cost you nothing.
  • A rebuilt example, where the names, the dates and the amounts have been replaced before uploading.
  • A login, a password, a key, an access token.
  • Health data, an identity document, bank details.
  • A document covered by professional confidentiality, a contract under a non-disclosure clause, a human resources file.
  • The work of a minor, handed in at school.
  • Content whose upload commits somebody who has asked you for nothing.

The habit to keep

Replace before sending. Most professional uses of a model do not need the real names, the real dates or the real amounts: they need the shape of the problem. A file rewritten with stand-in data gets the same help and pours nothing into the pool. It is the only habit in this chapter that depends on no service’s policy.

Going further

The mechanism that makes a model answer even when it knows nothing is covered in the chapter why it makes things up. The make-up of the data and its effects are covered in the chapter where the biases come from. The situations where the tool becomes a trap are gathered in when not to use it. A shorter read on the same subject is published on the blog: what becomes of an uploaded file. The contents of the course are on the understand page.

Sources