A fabricated voice
A voice is fabricated by stacking harmonics, then making them resonate like a vocal tract. This demonstration fabricates the same sound twice, once roughly and once carefully, and lets you compare what can be heard with what can be seen on its trace.
Everything is calculated in your browser. Both sounds are fabricated on this device, their spectrum is calculated here, and no sound is sent to or received from any server. Open your browser’s Network tab and use the demonstration: no request goes out. The two “Listen” buttons play a sound: turn the volume up if you want to hear it, or read the traces directly, since they are captioned to be read without it.
One sound, fabricated twice
The two sounds on this page start from the same point: the same pitch, the same seed, the same duration of one second. One is left raw, the other gets an extra treatment, described below in the section “The method”. Nothing else tells them apart.
Start the fabrication. Both sounds, their two spectrograms and the measured gap between harmonics and background will appear here, each with a “Listen” button.
What the demonstration shows
A train of harmonics, bare or shaped
Both sounds start from the same move: a series of harmonics tuned to the same pitch, the same seed, the same duration. The raw version stops there. The careful version does two more things: it shapes the amplitude of each harmonic through three fixed resonances, which give a vowel its colour, and it adds a micro-variation of pitch and amplitude from one period to the next, plus a low-level breath. Nothing else separates them.
On the trace, the difference can be seen before it can be heard: the raw version shows thin, perfectly sharp bands, with almost no energy between them. The careful version shows wider bands, concentrated at the three resonances, with a slightly greyed background between the harmonics: that is the breath.
A gap that vanishes with next to nothing
That gap, between the level of the harmonics and the level of the background between them, is measured in decibels on a single frame of sixty-four milliseconds taken from the middle of the sound. The values recorded at the default pitch, one hundred and thirty hertz, are these.
| Line noise added | Raw version | Careful version |
|---|---|---|
| none | 49.7 | 11.7 |
| 1% | 30.3 | 11.5 |
| 2% | 24.6 | 11.3 |
| 5% | 17.1 | 10.6 |
| 10% | 11.8 | 9.8 |
| 20% | 8.4 | 8.4 |
With no noise added at all, the gap in the raw version is nearly four times larger than the one in the careful version: a bare train of harmonics, with no breath and no irregularity, is sharper than any real voice ever is. Add the slightest line noise, and that excess sharpness collapses: at ten per cent, the gap in the raw version has already been divided by more than four. At twenty per cent, the two gaps are exactly equal: the measurement no longer tells the raw version from the careful one.
What makes a sound too sharp to be real is exactly what an ordinary telephone line wipes out first.
The method
What each version calculates, for anyone who wants to redo it without this page.
- Fabricate a train of harmonics tuned to the chosen pitch: for each whole rank, a sine wave at that rank times the fundamental frequency, with an amplitude of one over that rank. That is the shape of a simplified glottal source, common to both versions.
- Raw version: stop there. A very short amplitude envelope at both edges only avoids the click at the start and the end: nothing else is added.
- Careful version: shape the amplitude of each harmonic through three fixed resonances, the usual ones for the chosen vowel, not measured on any real person. Vary the pitch and the amplitude from one period to the next, by a small fraction drawn from the seed. Mix in a breath: low-level noise, drawn from the same seed.
- Bring both sounds to the same peak amplitude, so that the comparison never rests on a simple difference in volume.
- If the setting calls for it, add the same line noise afterwards to both sounds: it is what an ordinary transmission adds anyway, whether there is fabrication or not.
- Measure: take a frame of 1024 samples from the centre of the sound, calculate its spectrum by Fourier transform, record the level of each harmonic, record the mean level between them, and subtract.
Change the pitch or the vowel in the demonstration, and the resonances move with them: the measurement follows the sound, it never copies out a figure written in advance.
What this demonstration does not do
- It shows that a speech sound is fabricated in two stages: a source that vibrates, a tract that resonates, and that the second stage changes everything that makes a sound alive.
- It shows a measurable gap between the harmonics and the background of the spectrum, present in both versions.
- It shows from what line noise on that gap stops telling the two versions apart.
- It detects nothing. This page fabricates both sounds itself: it knows which is which, it does not guess.
- It judges no recording you might bring to it, and it accepts no file. A tool claiming to settle the matter on a real voice would be lying.
- It does not reproduce a real voice cloning system. Those learn their settings from a recording of a real person, whereas here the three resonances are set by hand, on a synthetic sound.
What research has measured
Three published results place what this page lets you feel in a minute.
Cloning a voice now takes very little material: a research paper describes a system, named VALL-E, able to synthesise personalised speech judged of high quality from a recording of only three seconds of a person it has never heard. Three seconds is less than the time it takes to say your own first name.
The clues that remain lie in signal-processing artefacts identifiable in fabricated audio, and research on detecting them relies precisely on that kind of trace to build its reference datasets. But the systems charged with spotting them still lack robustness in real acoustic conditions, and generalise poorly to generation methods they have not seen during training.
None of these sources measures the human ear directly against a recent cloned voice: that is therefore not a claim this page can make on their behalf. What they do allow anyone to say is more modest, and just as useful: a system trained specially to spot a fabricated voice already struggles to generalise to a new method, and an excess of spectral sharpness, where it exists, is exactly the kind of trace that ordinary line noise wipes out before it can serve anyone, human or detector.
A familiar voice, on the phone
In January 2024, a United States public authority, the Federal Communications Commission, documented and legally penalised a real fraud campaign: an automated call, sent out in New Hampshire ahead of the state’s presidential primary, carried a voice cloned by AI imitating President Joe Biden, and asked voters to stay at home and save their vote by skipping that primary. Nothing in the sound warned anyone that it had been fabricated.
A telephone call already compresses the sound, adds line noise to it, and leaves the ear only a narrow band of frequencies. That is exactly the combination this page has just fabricated with its own line noise setting: the conditions where a clue is most likely to disappear are those of an ordinary call, not those of a recording studio.
The habit to keep
Faced with a familiar voice asking, by telephone, for something urgent, a transfer, a code, an address, do not try to judge the voice: this page has just shown how unreliable that is. Hang up, and call back yourself on the number you already have, not the one that has just called. If the call has to carry on straight away, ask a question whose answer is written nowhere in public. A family or a team can also agree in advance on a word that only the real person knows, for the situations where urgency makes everything else hard to check. The habit costs nothing: the documented fraud above, on the other hand, cost dearly those who trusted their ears.
Going further
The same work on an image is in the chapter a fabricated image. The two ways of marking content at the source are in the chapter watermarks. The situations where the tool becomes a trap are gathered in when not to use it. All the demonstrations on the site are gathered on the demonstrations page, and the contents of the course on the understand page.
Sources
- WaveFake: A Data Set to Facilitate Audio Deepfake Detection Joel Frank, Lea Schönherr, 2021. Establishes that research on detecting synthetic voices relies on signal-processing artefacts identifiable in generated audio, and provides a reference dataset for studying them.
- ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild Xuechen Liu, Xin Wang, Md Sahidullah, Jose Patino, Hector Delgado, Tomi Kinnunen, Massimiliano Todisco, Junichi Yamagishi, Nicholas Evans, Andreas Nautsch, Kong Aik Lee, 2021. Establishes that systems for detecting synthetic speech still lack robustness in real acoustic environments, and generalise poorly to generation methods not seen during training.
- FCC Makes AI-Generated Voices in Robocalls Illegal (adopted statement: FCC 24-17) Federal Communications Commission (FCC), 2024. Establishes that a United States public authority documented and legally penalised a real fraud campaign using a voice cloned by AI imitating a political figure to reach voters ahead of an election.
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, Furu Wei, 2023. Establishes that the VALL-E system synthesises personalised speech judged of high quality from a recording of only three seconds of a person it has never heard.