This is an official ODERSA website. Here’s how you know

The official domain

The address of this site ends in odersa.org. Every service the association runs sits on a subdomain of odersa.org and nowhere else. If the address in your browser’s address bar does not end in odersa.org, this site is not ours.

Free, and no account

Everything is open straight away. No sign-up, no account, no password, no subscription, no advertising. Nothing is held back for those who pay, because there is nothing to pay for.

No data collected

This site does not follow you: no tracker, no tracking cookie, no measurement tool built into these pages, and nothing measured on your device. Our host counts requests in aggregate, as any server that answers does: a total, never a profile. You do not have to take our word for it: open your browser’s developer tools, go to the Network tab, and reload the page. You will see the full list of what the site asks for. Everything comes from odersa.org, nothing goes anywhere else.

Free to reuse

The content is published under the CC BY 4.0 licence. You may copy it, translate it, print it and pass it on, for your classes as much as for the people around you, on one condition only: credit ODERSA.

A fabricated voice

A voice is fabricated by stacking harmonics, then making them resonate like a vocal tract. This demonstration fabricates the same sound twice, once roughly and once carefully, and lets you compare what can be heard with what can be seen on its trace.

Several microphones on stands around a drum kit, in a recording studio
Recording a voice used to need a studio. Copying one now needs a few seconds of sound. Photograph: dan paluska, file page, licence CC BY 2.0 (licence text).
Information

Everything is calculated in your browser. Both sounds are fabricated on this device, their spectrum is calculated here, and no sound is sent to or received from any server. Open your browser’s Network tab and use the demonstration: no request goes out. The two “Listen” buttons play a sound: turn the volume up if you want to hear it, or read the traces directly, since they are captioned to be read without it.

One sound, fabricated twice

The two sounds on this page start from the same point: the same pitch, the same seed, the same duration of one second. One is left raw, the other gets an extra treatment, described below in the section “The method”. Nothing else tells them apart.

A word or a number. It sets the micro-variation of pitch, the grain of the breath and the line noise: the same seed gives exactly the same sound and the same trace, on any device.
Three fixed resonances give each vowel its colour. They change with this choice.
130 Hz
The fundamental frequency, in hertz. A low voice starts around 90, a high voice climbs towards 240.
0%
Added afterwards to both sounds, to imitate a telephone line or a compressed file. The most instructive setting on the page.

Start the fabrication. Both sounds, their two spectrograms and the measured gap between harmonics and background will appear here, each with a “Listen” button.

What the demonstration shows

A train of harmonics, bare or shaped

Both sounds start from the same move: a series of harmonics tuned to the same pitch, the same seed, the same duration. The raw version stops there. The careful version does two more things: it shapes the amplitude of each harmonic through three fixed resonances, which give a vowel its colour, and it adds a micro-variation of pitch and amplitude from one period to the next, plus a low-level breath. Nothing else separates them.

On the trace, the difference can be seen before it can be heard: the raw version shows thin, perfectly sharp bands, with almost no energy between them. The careful version shows wider bands, concentrated at the three resonances, with a slightly greyed background between the harmonics: that is the breath.

A gap that vanishes with next to nothing

That gap, between the level of the harmonics and the level of the background between them, is measured in decibels on a single frame of sixty-four milliseconds taken from the middle of the sound. The values recorded at the default pitch, one hundred and thirty hertz, are these.

Measured gap between the level of the harmonics and the level of the background of the spectrum, in decibels, according to the line noise added after fabrication. Seed “ODERSA”, pitch 130 Hz, vowel [a].
Line noise addedRaw versionCareful version
none49.711.7
1%30.311.5
2%24.611.3
5%17.110.6
10%11.89.8
20%8.48.4

With no noise added at all, the gap in the raw version is nearly four times larger than the one in the careful version: a bare train of harmonics, with no breath and no irregularity, is sharper than any real voice ever is. Add the slightest line noise, and that excess sharpness collapses: at ten per cent, the gap in the raw version has already been divided by more than four. At twenty per cent, the two gaps are exactly equal: the measurement no longer tells the raw version from the careful one.

What makes a sound too sharp to be real is exactly what an ordinary telephone line wipes out first.

The method

What each version calculates, for anyone who wants to redo it without this page.

  1. Fabricate a train of harmonics tuned to the chosen pitch: for each whole rank, a sine wave at that rank times the fundamental frequency, with an amplitude of one over that rank. That is the shape of a simplified glottal source, common to both versions.
  2. Raw version: stop there. A very short amplitude envelope at both edges only avoids the click at the start and the end: nothing else is added.
  3. Careful version: shape the amplitude of each harmonic through three fixed resonances, the usual ones for the chosen vowel, not measured on any real person. Vary the pitch and the amplitude from one period to the next, by a small fraction drawn from the seed. Mix in a breath: low-level noise, drawn from the same seed.
  4. Bring both sounds to the same peak amplitude, so that the comparison never rests on a simple difference in volume.
  5. If the setting calls for it, add the same line noise afterwards to both sounds: it is what an ordinary transmission adds anyway, whether there is fabrication or not.
  6. Measure: take a frame of 1024 samples from the centre of the sound, calculate its spectrum by Fourier transform, record the level of each harmonic, record the mean level between them, and subtract.

Change the pitch or the vowel in the demonstration, and the resonances move with them: the measurement follows the sound, it never copies out a figure written in advance.

What this demonstration does not do

  • It shows that a speech sound is fabricated in two stages: a source that vibrates, a tract that resonates, and that the second stage changes everything that makes a sound alive.
  • It shows a measurable gap between the harmonics and the background of the spectrum, present in both versions.
  • It shows from what line noise on that gap stops telling the two versions apart.
  • It detects nothing. This page fabricates both sounds itself: it knows which is which, it does not guess.
  • It judges no recording you might bring to it, and it accepts no file. A tool claiming to settle the matter on a real voice would be lying.
  • It does not reproduce a real voice cloning system. Those learn their settings from a recording of a real person, whereas here the three resonances are set by hand, on a synthetic sound.

What research has measured

Three published results place what this page lets you feel in a minute.

Cloning a voice now takes very little material: a research paper describes a system, named VALL-E, able to synthesise personalised speech judged of high quality from a recording of only three seconds of a person it has never heard. Three seconds is less than the time it takes to say your own first name.

The clues that remain lie in signal-processing artefacts identifiable in fabricated audio, and research on detecting them relies precisely on that kind of trace to build its reference datasets. But the systems charged with spotting them still lack robustness in real acoustic conditions, and generalise poorly to generation methods they have not seen during training.

None of these sources measures the human ear directly against a recent cloned voice: that is therefore not a claim this page can make on their behalf. What they do allow anyone to say is more modest, and just as useful: a system trained specially to spot a fabricated voice already struggles to generalise to a new method, and an excess of spectral sharpness, where it exists, is exactly the kind of trace that ordinary line noise wipes out before it can serve anyone, human or detector.

A familiar voice, on the phone

In January 2024, a United States public authority, the Federal Communications Commission, documented and legally penalised a real fraud campaign: an automated call, sent out in New Hampshire ahead of the state’s presidential primary, carried a voice cloned by AI imitating President Joe Biden, and asked voters to stay at home and save their vote by skipping that primary. Nothing in the sound warned anyone that it had been fabricated.

A telephone call already compresses the sound, adds line noise to it, and leaves the ear only a narrow band of frequencies. That is exactly the combination this page has just fabricated with its own line noise setting: the conditions where a clue is most likely to disappear are those of an ordinary call, not those of a recording studio.

The habit to keep

Faced with a familiar voice asking, by telephone, for something urgent, a transfer, a code, an address, do not try to judge the voice: this page has just shown how unreliable that is. Hang up, and call back yourself on the number you already have, not the one that has just called. If the call has to carry on straight away, ask a question whose answer is written nowhere in public. A family or a team can also agree in advance on a word that only the real person knows, for the situations where urgency makes everything else hard to check. The habit costs nothing: the documented fraud above, on the other hand, cost dearly those who trusted their ears.

Going further

The same work on an image is in the chapter a fabricated image. The two ways of marking content at the source are in the chapter watermarks. The situations where the tool becomes a trap are gathered in when not to use it. All the demonstrations on the site are gathered on the demonstrations page, and the contents of the course on the understand page.

Sources