A fabricated image
An image made by a machine often carries the trace of the operations that produced it, and that trace can be measured. This demonstration fabricates two images on your device, shows where the trace appears, and shows from what point on it becomes unreadable.
Everything is calculated in your browser. Both images are drawn on this device, their spectra are calculated here, and no image is downloaded from anywhere. Open your browser’s Network tab and use the demonstration: no request goes out.
Two images, the same content, one operation of difference
The two images below come out of the same starting field. One is rendered as it is. The other goes through a bottleneck: it is shrunk, then enlarged again by repeating each pixel, which is the upsampling operation that image generators chain together to go from a small map of values to a large image. Nothing else separates them.
Start the fabrication. Both images, their two spectra and the measured figures will appear here.
What the measurement shows
The trace is a gap, not a peak
The common idea is that a fabricated image shows peaks in its spectrum. What this demonstration measures is the opposite, and it is more interesting. At the frequencies of the upsampling grid, the fabricated image consistently has less energy than its surroundings, by around eleven decibels. Those frequencies are exactly the ones that the repetition kernel cancels: by repeating each pixel a certain number of times, you apply a filter whose zeros fall at multiples of that number. The spectrum therefore keeps the dip left by the tool that produced it.
On the control image, with no noise added, the same measurement gives less than one decibel of difference, one way or the other, because there is no reason for a texture to favour or avoid those particular frequencies. The difference between the two images is therefore not an impression: it is a number, and it can be reproduced.
It disappears under next to nothing
Push the noise slider and watch the measurement melt away. The values recorded on the default seed, at factor eight, are these.
| Noise added | Control image | Fabricated image |
|---|---|---|
| none | +0.2 | -11.0 |
| 2 levels | -0.5 | -4.6 |
| 5 levels | -0.8 | -1.6 |
| 10 levels | -1.2 | -0.5 |
| 20 levels | -0.7 | +0.2 |
| 40 levels | -0.4 | +0.8 |
At five levels of grey of noise, that is two per cent of the scale, the trace has already melted away: from eleven decibels it drops to less than two, just double what the control shows on its own. At ten levels it no longer exists: the gap measured on the fabricated image falls below the one on the control, and nothing tells them apart any more. Noise of that size is invisible to the eye, and it happens all on its own: a recompression, a screenshot, a resize, a retouching filter do as much without anyone meaning it.
The trace is beyond dispute and it survives next to nothing. Both halves of that sentence count equally.
The method
The whole calculation, for anyone who wants to redo it without this page.
- Fabricate the starting field. Five layers of smoothed noise, finer and finer, added together with an amplitude that halves at each layer. The result is a texture whose energy falls off with frequency, like that of a photograph. The draw is seeded from the seed shown, so it is reproducible.
- Fabricate the second image. Average the field in square blocks of the chosen side, which gives a small image, then enlarge it again by repeating each pixel that many times. This is nearest-neighbour upsampling.
- Add the noise, if there is any, at the same amplitude on both images.
- Convert each image to luminance, remove its mean, then calculate its two-dimensional Fourier transform. Removing the mean is essential: otherwise the constant component swamps everything and the spectrum is nothing but a point in the middle.
- Recentre the spectrum, take the amplitude of each frequency, and convert it into decibels relative to the maximum.
- Measure. Average the decibels at the frequencies whose two coordinates are both multiples of the grid step, then average the decibels everywhere else, and subtract. That is the grid gap. A value close to zero means those frequencies are nothing special. A clearly negative value means a repetition kernel has been through.
The grid step is the side of the image divided by the upsampling factor. Change the factor in the demonstration, and the step changes with it: the measurement goes looking for the trace somewhere else, and it finds it. That is what shows the measurement really follows the operation, and does not find what it wants to find.
What this demonstration does not do
- It shows that an operation of fabrication leaves a measurable and reproducible trace in the frequency domain.
- It shows that this trace depends on the setting of the operation, and moves with it.
- It shows at what level of disturbance the trace stops being readable.
- It detects nothing. This page fabricates both images itself: it knows which is which, it does not guess, and it returns no confidence score.
- It judges no image you might bring to it, and it accepts no file. A tool claiming to settle the matter on a real image would be lying.
- It does not reproduce a real image generator. Upsampling is one of its operations, not all of them, and recent architectures deliberately soften this signature.
What research has measured
Three published results frame this demonstration, and a fourth gives its limit.
Images produced by generative networks show systematic artefacts visible in the frequency domain, caused by the upsampling operations common to those architectures. The upsampling component of adversarial networks leaves a characteristic and reproducible spectral fingerprint, usable even without having seen that particular model during training. That work says the same thing as the measurement above: the tool signs its passage.
The limit is published too, and it is clear. A peer-reviewed paper establishes that watermarks invisible at pixel level can be removed by a regeneration attack, which adds noise and then reconstructs the image, and that this family of attacks preserves visual quality. Its authors conclude that invisible watermarks must give way to marks that preserve the meaning of the image. What the noise slider on this page does in three seconds is the elementary version of that attack.
One last fact closes the discussion about the human eye: observers can no longer reliably tell a synthetic face from a real one, and even judge synthetic faces on average more trustworthy. Trusting your own impression is therefore not a fallback method.
The habit to keep
Faced with an image whose origin matters, do not look for the clue in the image. Look for the source: who published it first, on what date, and what the provenance metadata say when they exist. That route is covered in the chapter watermarks. A technical clue serves to document a doubt, never to conclude on its own.
Going further
The same work on a sound signal is in the chapter a fabricated voice. The two ways of marking content are in the chapter watermarks. The situations where the tool becomes a trap are gathered in when not to use it. All the demonstrations on the site are gathered on the demonstrations page, and the contents of the course on the understand page.
Sources
- Leveraging Frequency Analysis for Deep Fake Image Recognition Joel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer, Dorothea Kolossa, Thorsten Holz, 2020. Establishes that images generated by generative networks show systematic artefacts visible in the frequency domain, caused by the upsampling operations common to those architectures.
- Detecting and Simulating Artifacts in GAN Fake Images Xu Zhang, Svebor Karaman, Shih-Fu Chang, 2019. Establishes that the upsampling component used by adversarial networks leaves a characteristic and reproducible spectral fingerprint, usable to detect a generated image even without having seen that particular model during training.
- Invisible Image Watermarks Are Provably Removable Using Generative AI Xuandong Zhao, Kexun Zhang, Zihao Su, Saastha Vasan, Ilya Grishchenko, Christopher Kruegel, Giovanni Vigna, Yu-Xiang Wang, Lei Li, NeurIPS 2024. Establishes that a regeneration attack, which adds noise to an image and then reconstructs it, removes watermarks invisible at pixel level while preserving visual quality, and that the authors accordingly recommend moving to marks that preserve the meaning of the image.
- AI-synthesized faces are indistinguishable from real faces and more trustworthy Sophie J. Nightingale, Hany Farid, 2022. Establishes that human observers can no longer reliably tell a synthetic face from a real one, and even judge synthetic faces on average more trustworthy.