It does not know that it does not know
A language model does not lie and does not know that it is wrong: it always produces the most probable word, whether that word is true or not.
A language model can state a false fact with exactly the same assurance as a true one. It is not a lie, and it is not an error the model could avoid by paying more attention either. The reason lies in the way a text is made, and in the way these models are evaluated during their training.
What a model does, exactly
A language model does not consult a database of facts to answer. It produces, token after token, the most probable word given everything that comes before. That production carries on in the same way whether the sentence being written is accurate or not: nothing in the calculation stops to check a fact against the real world. A false answer is written with the same fluency as a right one, because it goes through the same mechanism.
An exam where guessing pays more than holding back
One question remains: why does training not correct this tendency towards confidence? A study by Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala and Edwin Zhang, published in 2025, answers with a precise comparison. The procedures that evaluate a model during its training often work like a multiple-choice exam marked the old way: a right answer earns a point, a box left blank earns none, and a wrong answer costs none either. On that marking scheme, guessing always remains the best strategy, even with no knowledge at all. A model trained and then marked by similar rules learns the same lesson: answering, rather than admitting uncertainty, gains ground on the scores that count.
Like pupils facing a hard question, models sometimes guess rather than admit their uncertainty.
The authors do not stop at a diagnosis. They put forward an answer as simple to state as it is hard to generalise: changing the marking scheme of evaluations so that admitting uncertainty costs less than a wrong answer. As long as that is not the norm on the scoreboards that rank models, guessing remains the strategy that pays most in training.
Why it cannot be seen from the inside
Nothing in the text produced carries any trace of that uncertainty. A model has no separate module assessing its own reliability before writing a sentence: no internal boundary separates “I know” from “I am guessing”. That absence, documented by the survey by Lei Huang and his co-authors on the whole set of causes and forms of hallucination in large language models, explains why the error is so hard to spot from outside: it announces itself by no sign in the tone, the length or the structure of the sentence.
An unsettled debate: is hallucination unavoidable?
One point goes further, and it remains contested. In 2024, Ziwei Xu, Sanjay Jain and Mohan Kankanhalli published a formal demonstration, based on computability theory: a language model, as a computable function, cannot learn every possible computable function, and will therefore always produce, over an infinite set of possible questions, at least one false answer. The authors present that result as a structural limit holding for any model of this kind. It is theoretical work, not an observation of use: the authors themselves note that this formal world remains simpler than the real one.
In 2025, Atsushi Suzuki, Yulan He, Feng Tian and Zhongyuan Wang answered without contradicting the demonstration: they accept it, but dispute its reach. In their view, even if hallucination cannot disappear entirely over an infinite set of cases, its probability can always be reduced, to the point of becoming negligible, by improving the training data and the methods. A theoretical impossibility result, obtained by an argument about a world of computable functions, would therefore say nothing about the problems genuinely met in use. The disagreement is less about the mathematics, which nobody questions, than about what it allows anyone to conclude for a model genuinely in use.
The case for unavoidability
A language model is a computable function. It cannot learn every possible computable function. It will therefore always hallucinate, on at least one question, however good its training (Xu, Jain, Kankanhalli, 2024).
The 2025 answer
The demonstration holds, but only over an infinite set of cases. In practice, the probability of a hallucination goes down with better data, to the point of becoming negligible (Suzuki, He, Tian, Wang, 2025).
What the two texts share, despite their disagreement, comes back to the starting observation: neither theory nor practice allows a model to tell, at the moment it writes, an uncertain answer from a false one stated with the same assurance. The chapter of the course devoted to why a model makes things up sets out the mechanism behind that observation.
Sources
- Why Language Models Hallucinate Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang, 2025.
- A Survey on Hallucination in Large Language Models Lei Huang et al., 2023.
- Hallucination is Inevitable: An Innate Limitation of Large Language Models Ziwei Xu, Sanjay Jain, Mohan Kankanhalli, 2024. Theoretical work whose practical reach is contested: see the next source.
- Hallucinations are inevitable but can be made statistically negligible Atsushi Suzuki, Yulan He, Feng Tian, Zhongyuan Wang, 2025.
This article is published under the CC BY 4.0 licence: copy it, translate it, republish it, crediting ODERSA.