AI detectors at school
A detector of AI-generated text promises a simple answer. The published research on these tools measures mistakes, and those mistakes do not fall at random on everyone alike.
A detector of AI-generated text promises a simple answer to a simple question: was this text written by a person or by a machine? The published research on these tools says something else: they get it wrong, and their mistakes do not fall at random on everyone alike. Schools are nevertheless already building these tools into their marking procedures, often without knowing what a score really means.
What a detector really measures
A detector does not recognise a signature left by a machine, like a fingerprint. It compares the statistical regularity of a text, the predictability of its successive words, with what generally characterises a text produced by a language model. The result is a probability score, not proof: the same tool can change its verdict on the same text slightly reworded.
Two mistakes, and they do not cost in the same place
A detector can get it wrong in two ways: wrongly accusing a human text of being generated, or letting a generated text through as if it were human. A 2023 study, led by Weixin Liang and his co-authors, measures the first mistake on a precise point: essays written by non-native speakers of English are systematically classified as AI-generated, while essays by native speakers, on a comparable subject, are correctly recognised as human. The cause lies in vocabulary: less varied, more predictable turns of phrase, which statistically resemble what a model produces, without a model ever having touched them.
The widest independent evaluation on the subject, carried out in 2023 by Debora Weber-Wulff and 7 co-authors, tests 14 tools: 12 open to the public and 2 commercial ones already used by schools. Conclusion: not one of the 14 is both reliable and accurate. The dominant bias runs the other way from the previous study: these tools tend to classify a text as human rather than to detect a generated text correctly, and a slight alteration of the text is often enough to fool them.
A maker withdraws its own tool
OpenAI launched a classifier in January 2023, meant to tell a human text from one produced by its own models. Six months later, the company withdrew it. On 20 July 2023, it updated the tool’s presentation page to announce its withdrawal on account of its low rate of accuracy, a fact relayed on 25 July by the site TechCrunch, which had itself tested the tool at launch.
The maker that knows its own model best did not manage to turn it into a reliable detector.
If the company best placed to know its own model did not manage it, the difficulty is not a matter of insufficient effort on one occasion: it lies in the very nature of the task.
Who pays for the mistake
The two directions of the mistake do not weigh on the same people. A detector that is too suspicious punishes, first of all, a pupil writing in a language that is not their own, or who simply writes in a plain and even way. A detector that is too trusting lets a generated text through without flagging it to anyone. In both cases, the score displayed has the appearance of an objective measurement, when none of the public studies cited here validates it as reliable at the scale of a school. When a decision rests on that score, the burden of proof often shifts to the accused pupil, who then has to account for a piece of work that no validated tool can, for its part, confirm or rule out with certainty.
What a detection score makes it possible to do, and what it does not:
- Open a conversation with a pupil about how a piece of work was produced.
- Serve as one clue among others, to be cross-checked against knowledge of the pupil’s usual work.
- Amount, on its own, to proof of cheating: not one of the 14 tools tested in 2023 was validated as reliable at that level.
- Treat every pupil in the same way: the documented mistake hits more often those writing in a language that is not their own.
The paths of The AI Manual designed for those who teach come back to how to respond to a suspicion, without turning a statistical score into a verdict.
Sources
- GPT detectors are biased against non-native English writers Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou, 2023.
- OpenAI scuttles AI-written text detector over “low rate of accuracy” Devin Coldewey, TechCrunch, 2023.
- Testing of detection tools for AI-generated text Debora Weber-Wulff et al., 2023.
This article is published under the CC BY 4.0 licence: copy it, translate it, republish it, crediting ODERSA.