This is an official ODERSA website. Here’s how you know

The official domain

The address of this site ends in odersa.org. Every service the association runs sits on a subdomain of odersa.org and nowhere else. If the address in your browser’s address bar does not end in odersa.org, this site is not ours.

Free, and no account

Everything is open straight away. No sign-up, no account, no password, no subscription, no advertising. Nothing is held back for those who pay, because there is nothing to pay for.

No data collected

This site does not follow you: no tracker, no tracking cookie, no measurement tool built into these pages, and nothing measured on your device. Our host counts requests in aggregate, as any server that answers does: a total, never a profile. You do not have to take our word for it: open your browser’s developer tools, go to the Network tab, and reload the page. You will see the full list of what the site asks for. Everything comes from odersa.org, nothing goes anywhere else.

Free to reuse

The content is published under the CC BY 4.0 licence. You may copy it, translate it, print it and pass it on, for your classes as much as for the people around you, on one condition only: credit ODERSA.

Watermarks

Two mechanisms go by the name of “watermark”, and they are endlessly confused. This chapter fabricates one of each kind before your eyes, in a text and then in an image, and destroys the second in front of you to show its real limit.

Information

Everything is calculated in your browser. The text is generated by a toy model learnt on this device, the image is drawn on a local canvas, and the message it hides is never sent anywhere. Open your browser’s Network tab and use the demonstration: no request goes out.

Two mechanisms people confuse

Two very different mechanisms are called a “watermark”, and that confusion comes at a price. The first declares the origin of a file in its metadata, like a label stuck on it. The second quietly alters the content itself, like a texture woven into the cloth. A file can carry one, the other, both, or neither.

This chapter fabricates the second mechanism before your eyes, in a text and then in an image, because it can be calculated. The first cannot be calculated: it is declared, or it is not. Then comes the comparison of the two, with what each survives and what neither of them replaces.

A green-list watermark, placed on a text

At every word it generates, this site’s toy model works out a list of possible candidates. A secret key, combined with the preceding word, marks one half of that list as green. The draw is then nudged slightly towards that half. Set the key and the strength of that nudge, then generate a text: each word produced appears with the label the key gives it at that moment.

A word or a number. The same key has to be used again to find the watermark afterwards. A real service keeps it secret; here it stays visible so that the calculation can be checked.
3.0
At zero, no nudge at all: the draw ignores the green list. The higher the strength, the more it favours that list, at the price of a more constrained text.
The model continues this opening, as in the chapter the next word. The corpus being a French novel, so is the opening.

The generated text will appear here, word by word, with the colour the key gives it.

The text above can be measured with the same key, and so can any other text: paste one in, or change the one that has just appeared, then measure. The result is never watermarked or not watermarked: it is a proportion, to be compared with what it would be without a watermark.

Pre-filled with the text generated above. Replace it with any other text to measure that one in turn, with the same key entered above.

The proportion of words on the green list, and how to read it, will appear here.

A watermark in the least significant bits of an image

A digital image is a grid of numbers, one per colour channel and per pixel, from zero to two hundred and fifty-five. Changing the least significant bit of just one of those numbers alters its value by at most one: nothing anyone can see. This demonstration draws an image, hides a message in exactly those bits, finds it again, then lets you destroy it.

It determines the background drawing. The same seed always gives back the same image, on any device.
A few dozen characters fit easily into an image of this size.

The fabricated image, and the message found on it, will appear here.

Now choose an everyday operation, one of those that sending by messaging app or sharing on a network applies all by itself, and see whether the message survives.

Four choices, none of them designed to wipe out a watermark: they are ordinary moves.

The image after the chosen operation, and what is left of the message, will appear here.

The method

The green-list watermark, step by step

  1. For the current word, obtain the list of possible candidates and their probability, exactly as in the chapter the next word.
  2. Combine the secret key with the immediately preceding word, then draw from that a number between zero and one through a hash function. Below one half, the candidate is green; above, it is red. Half the vocabulary is green at every step, but never the same half: it changes with the preceding word.
  3. Add the chosen strength to the logit of every green candidate, before the temperature and the draw: that is the bias the toy model accepts at this very stage.
  4. Draw normally from that corrected distribution. A green word now has a better chance of being chosen, without any word becoming impossible; and when all the candidates at a given step share the same colour, the bias changes nothing, there being no choice to separate.
  5. Repeat for the next word: the preceding word has changed, so the green list has too.
  6. To measure a text, redo the calculation of step two on each of its words, with the same key, and count the share of green words. Under the key that served to generate it, that share clearly exceeds one half; on any other text, or under another key, it stays close to one half.
  7. Compare that measured share with the expected half through a standardised difference: the bigger the difference, the rarer the observed proportion would be in a text that had not received this watermark. That number never becomes a verdict.

Hiding a message in the least significant bits, step by step

  1. Draw the image and read its pixels: for each one, three numbers from zero to two hundred and fifty-five, one per colour channel.
  2. Write the length of the message, in bytes, into the first thirty-two bits available.
  3. Write each byte of the message after it, eight bits each.
  4. For each bit to be written, take the next number in the grid of colour channels, wipe its least significant bit and put the message bit there. The number changes by at most one, on a scale of two hundred and fifty-five: no eye sees the difference.
  5. To find the message again, read the same bits in the same order, rebuild the length and then the bytes, and convert them into text.
  6. An operation that recalculates the pixels, a recompression or a resizing, recalculates their least significant bit too, with no regard for what it was carrying. The message survives only if that value stays identical, to the bit.

Two ways of marking content

The two attempts above put the two mechanisms this chapter distinguishes in their proper place.

Declared provenance: readable, and fragile

A public technical standard, C2PA, says how to attach signed information about its origin and its history to a file: which device or which software produced it, what changes it has been through. An application that reads this standard displays that information in plain sight, under the name Content Credentials. It is a label, in the literal sense: it says where the file comes from, it touches nothing inside the file itself.

A label comes off. The standard itself acknowledges it: its provenance information can be separated from the file, which makes it losable through a simple screenshot or a re-saving of the file. Sending by a messaging app that recompresses an image, a screenshot that copies only the pixels displayed, an export to a format that ignores that block of metadata: none of those moves is an attack, and each is enough to wipe out the provenance without leaving any trace of the wiping.

That is where the most widespread misreading lies: an image with no declared provenance is not an image hiding something. It may well be an image whose label simply fell off along the way, which happens to very nearly all the images in circulation.

The statistical watermark: discreet, and more resistant, not invincible

A statistical watermark is not added alongside the content: it alters the content itself, in a tiny and spread-out way. In a text, by nudging the draw towards a list of words chosen by a key, as the demonstration above has just done. In an image, by altering pixels in a way that can be far more discreet than the least significant bits used here to stay explainable.

That difference in nature changes everything about resistance. A metadata label disappears in a single move, because it is never more than data beside the data. A statistical watermark has to be undone inside the content itself: copying a text out by hand changes nothing about its green list, and an image sometimes keeps a watermark through operations that would wipe out its provenance with no effort.

That resistance has two limits, and both have just been seen above. The first: a watermark assumes that the maker placed it at the moment of creation. Nothing obliges a service to do so, and nothing proves it after the fact if the content has circulated without it. The second: a watermark is not invincible either. A paper cited in the chapter a fabricated image shows that a regeneration attack, which adds noise to an image and then reconstructs it, removes watermarks invisible at pixel level while preserving visual quality. The recompression or the resizing you have just chosen above is the elementary version of it, applied to a different mechanism but to the same idea: recalculating the pixels wipes out whatever was hiding in them.

  • A statistical watermark alters the content itself, so it travels with copies that a metadata label does not follow.
  • Detecting it returns a measured proportion and a measured difference, never a word to settle the matter.
  • Watermarks of this family really are deployed and published, peer reviewed.
  • It does not place itself: without a maker choosing to apply it, there is nothing to detect.
  • It does not resist everything: a firm enough operation, even an ordinary one, does away with it.
  • It never says whether content is true or false, only whether it carries the mark of a given piece of fabrication.

What research has measured

A public technical standard defines how to attach verifiable, signed information about its origin and its history of changes to a file, and acknowledges itself that this information is separable from the file, and therefore losable through a screenshot or a re-saving, with nothing to signal the loss.

Inserting into a generated text a statistical watermark invisible to the eye, by slightly favouring certain words at every step, then detecting it by a statistical test: that is exactly the mechanism the demonstration above has just run. A mechanism of this family, built into the sampling of a real service, has been published and peer reviewed in the journal Nature: what this page simplifies for teaching purposes is deployed elsewhere on an entirely different scale.

The habit to keep

Faced with a file whose origin matters, two reflexes too often stand in for a judgement. The first: looking for a declared provenance and, failing to find one, concluding that the file is hiding something. The standard itself says the opposite: the absence of provenance is the normal state of very nearly all the files in circulation, not proof of anything. The second: asking a detection tool for a verdict, on a text or on an image. An independent evaluation of fourteen text detection tools, including tools used in education, concludes that none is both reliable and accurate; the maker of one of the most used withdrew its own tool six months after launch, citing a rate of accuracy too low to be useful.

The green-list watermark demonstration above applies that same caution to itself: its result is a proportion and a standardised difference, never a word to settle the matter. Faced with content whose origin really matters, the useful question stays the same: who published it first, on what date, and is it corroborated elsewhere. A technical clue documents a doubt, it never replaces one.

Going further

The mechanism this chapter sets aside, the one that spots the unintended trace of a fabrication rather than a watermark placed on purpose, is covered in the chapters a fabricated image and a fabricated voice. The draw that the text watermark nudges is explained in detail in the chapter the next word. The situations where none of these tools is enough are gathered in when not to use it. All the demonstrations on the site are gathered on the demonstrations page, and the contents of the course on the understand page.

Sources