Every large language model is, underneath the branding, a guessing machine, and the guessing is how it produces anything at all rather than a defect that better engineering will someday retire.

When a model produces a sentence it is not retrieving a fact and reciting it; it is predicting which token is most likely to come next, given everything before it. Most of the time that lines up with reality, because reality is what the training data mostly described. Sometimes it does not, and what comes out is fluent and confident and wrong: a case citation for an opinion nobody wrote, or a quotation that appears nowhere in the source it is attributed to, or a number with no study behind it. The research literature separates two failure modes that casual use runs together. Factuality errors conflict with the world; faithfulness errors conflict with the source the model was handed.3 Neither one announces itself in the output.

What is actually happening when a model makes something up?

OpenAI's own account of this is unusually plain. Pretraining is next-word prediction over enormous amounts of text, and the model sees only positive examples of fluent language; nothing in the data is labeled false. Spelling and parenthesis matching follow consistent patterns, so those errors disappear as models get larger. Arbitrary low-frequency facts do not follow patterns and cannot be predicted from them, and that is where the invention happens.2

The paper behind that post makes the same point formally. Hallucinations, it argues, originate as errors in binary classification: if incorrect statements cannot be distinguished from correct ones, some rate of error follows from statistical pressure rather than from anything mysterious.1 The paper is "Why Language Models Hallucinate," posted on 4 September 2025 by Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala and Edwin Zhang.1 OpenAI published a companion post the following day.2

Asked for the title of Kalai's own doctoral dissertation, a widely used chatbot produced three different answers, none of them correct. Asked for his birthday, it gave three different dates, all of them wrong.2 That is three confident errors on request, about one of the authors of the paper, and nothing in any of the outputs marks which of the three might be the real one.

Why does the guessing not stop?

The incentive half of the argument is the half I find persuasive, and I say that as a matter of professional judgment rather than machine learning.

Nearly every benchmark used to rank models grades an answer right or wrong, with no partial credit for declining to answer. A model that says it does not know scores zero, which is exactly what it scores for being flatly wrong, while a model that guesses has some chance of being marked correct. Across thousands of questions the guesser wins the scoreboard.12

On the SimpleQA evaluation reported in OpenAI's GPT-5 system card, one model abstained on 52 percent of questions, answered 22 percent correctly and got 26 percent wrong. An older model abstained on 1 percent, answered 24 percent correctly and got 75 percent wrong.2 The older model wins by two points on accuracy while being wrong nearly three times as often, and an accuracy-only leaderboard reports the two-point win and stays quiet about the rest.

None of that is a law of nature; it is a scoring rule. A scoreboard that pays for a correct answer and charges nothing for a wrong one is a bonus plan, and people build to the bonus plan. So do models. The paper's proposed fix follows directly from that reading: change the scoring of the benchmarks that already dominate the leaderboards, rather than publishing one more hallucination test to sit alongside them.12 A test nobody is ranked on does not change what gets built.

Can hallucination be eliminated, or only managed?

Two lines of formal work say the error itself cannot be removed. Ziwei Xu, Sanjay Jain and Mohan Kankanhalli define a formal world in which hallucination is an inconsistency between a computable language model and a computable ground truth function, then use results from learning theory to show that such models cannot learn all computable functions and will therefore hallucinate if used as general problem solvers.4 Sourav Banerjee, Ayushi Agarwal and Saloni Singla argue from undecidability results and Godel's first incompleteness theorem that every stage of the pipeline, from training data compilation through retrieval, intent classification and generation, carries a non-zero probability of producing hallucination.5

OpenAI disagrees in writing and by name. Its post lists "hallucinations are inevitable" as a claim and answers that they are not, because a model can abstain when it is uncertain.2 In the same list it concedes that accuracy will never reach 100 percent, because some real questions are unanswerable regardless of model size, search or reasoning.2

The two camps are not measuring the same object. The formal results bound error; OpenAI is talking about the confident assertion of it. Being wrong is not eliminable, but sounding certain while being wrong is a behaviour, and behaviours respond to how you grade them. That is why "solve hallucination" is the wrong project brief. The honest limitation on both formal papers is that they bound a mathematical object under a definition the authors chose for themselves, not a shipped product under a support contract.

Is the guessing ever the point?

The consultancy Thoughtworks published a piece in August 2025 by Abhishek Roy arguing that hallucination is the natural result of computing that runs on probability rather than strict logic, and therefore not a system failure at all, and that the work now is learning to use such a system safely rather than waiting for the behaviour to be patched out.6 Roy reaches for the image of a student who does not know the answer and writes around the topic instead of leaving the page blank.6 Two weeks later the OpenAI paper opened on the same image: like students facing hard exam questions, models guess when uncertain.1 When a consultancy and a research lab reach for the same analogy inside a fortnight, the analogy is doing real work.

The slogan version of this, that hallucination is the price of creativity, has more support than nothing and less than it usually gets. Zicong He, Boxuan Zhang and Lu Cheng built a framework to quantify hallucination and creativity across the decoding layers of a model and found a tradeoff that held across layer depth, model type and model size, with one layer at each size balancing the two best.7 That is a real result on a narrow definition of creativity that the authors set for themselves and said so.7

A separate group tested the assumption from the other end, by asking what actually happens to creative output when hallucination-reduction techniques are applied. Across three methods and several model families at sizes from one billion to seventy billion parameters, the effects ran in opposite directions: chain of verification increased divergent thinking, contrastive decoding suppressed it, and retrieval augmentation made little difference either way.8 Reducing hallucination does not uniformly cost you ideas. It depends on which reduction method you choose, which is more useful than the slogan and considerably less quotable.

The defensible claim is narrower than the one made from conference stages. A model that will only say what it can cite makes a worse brainstorming partner, and a model that will say anything makes a worse research assistant. Those are two settings of one dial, and the dial belongs to whoever is doing the work.

What does it cost when the answer has to be true?

Damien Charlotin maintains a public database of legal decisions in which a court or tribunal addressed a party's reliance on hallucinated material. Read on 24 August 2026, it listed 1,959 decisions, 1,344 of them from the United States. Lawyers accounted for 781 and self-represented litigants for 1,127. Fabricated material appeared in 1,630 entries and false quotations in 527.9 The database tracks decisions where a court engaged with the problem, not the wider universe of filings that carried fake citations and were never caught, so the real denominator is larger and unknown.9

The founding case is still the clearest statement of principle. In June 2023, Judge P. Kevin Castel sanctioned two attorneys and their firm $5,000, jointly and severally, for submitting non-existent judicial opinions with fake quotes and citations produced by ChatGPT, then standing by those opinions after the court questioned whether they existed.10 His framing has aged better than most commentary since: there is nothing inherently improper about using a reliable artificial intelligence tool for assistance, but existing rules impose a gatekeeping role on attorneys to ensure the accuracy of their filings.10 A federal judge wrote the position of this essay three years before I did.

The penalties have grown. In December 2025 a magistrate judge in the District of Oregon found that a plaintiff's three summary judgment briefs contained citations to fifteen non-existent cases along with fabricated quotations falsely attributed to eight legitimate authorities. The briefs were stricken without leave to refile, counsel was ordered to pay $15,500 to the clerk of the court, and the defendants were awarded their fees.11 The opinion declines to print the names of the fictitious cases, reasoning that doing so could amplify the lie that they exist.11 That is control design, and it costs nothing.

The Nebraska Supreme Court reached the sharper question in March 2026. Its opinion sets out a chart of twenty problematic citations from a single appellate brief: fictitious cases, real cases carrying fictitious quotations, real court rules carrying fictitious quotations, real cases cited for propositions they do not support. One fabricated authority, wearing the name of a real unpublished decision, was cited and quoted five times through the brief.12 Counsel denied using AI at all and attributed the errors to a cracked laptop screen and the filing of a wrong draft.12

Regardless of whether AI was used, the court held, the analysis is the same.12 It struck the brief, dismissed the appeal and referred counsel to the state disciplinary authority, quoting the observation that citing nonexistent case law or misrepresenting the holding of a case is making a false statement to a court, and that it does not matter if generative AI told you so.12 The duty did not move when the tool arrived, and that holds well past law.

It reaches transcription, for one, where the failure is quieter. Allison Koenecke, Anna Seo Gyeong Choi, Katelyn X. Mei, Hilke Schellmann and Mona Sloane evaluated OpenAI's Whisper and found that roughly 1 percent of audio transcriptions contained entire hallucinated phrases or sentences that existed nowhere in the underlying audio. Thirty-eight percent of those hallucinations carried explicit harms, including invented associations and implied false authority. The hallucinations fell disproportionately on speakers with longer non-vocal stretches, a common symptom of aphasia.13 A 1 percent rate reads as tolerable until you multiply it by volume and then ask who absorbs the error. In that study it landed on the speakers least able to correct the record.

Chief Justice John Roberts flagged the pattern early, writing in his 2023 year-end report on the federal judiciary that any use of AI requires caution and humility, and pointing directly at the fake-citation episode then in the news.14 Three years and roughly two thousand recorded decisions later, the advice has not been widely taken.

Does anyone have to tell you the error rate?

No, and the disclosure that does exist is getting thinner. The 2026 AI Index from Stanford's institute for human-centered artificial intelligence reports that almost all leading frontier developers publish results on capability benchmarks such as MMLU and SWE-bench, while reporting on responsible AI benchmarks remains sparse.15 The Foundation Model Transparency Index puts a number on the same trend: the average score rose from 37 to 58 between 2023 and 2024, then fell to 40 in 2025, with the widest gaps in disclosure of training data, compute and post-deployment impact.15

The same chapter describes a new accuracy benchmark on which hallucination rates across 26 leading models ranged from 22 percent to 94 percent, and on which accuracy for GPT-4o dropped from 98.2 percent to 64.4 percent and for DeepSeek R1 from above 90 percent to 14.4 percent.15 The finding underneath those numbers is about framing. Present a false statement as something another person believes and models handle it; present the same false statement as something the user believes and performance collapses.15 Documented incidents rose to 362 in 2025 from 233 the year before.15 One genuinely encouraging number sits in the same chapter: the share of surveyed businesses with no responsible AI policy in place fell from 24 percent to 11 percent, with knowledge gaps and budget named as the leading obstacles to doing more.15

Set that against deployment. The same report puts organizational adoption at 88 percent of surveyed organizations, with generative AI used in at least one business function at 70 percent.16 So adoption is near universal, disclosure of error rates is voluntary and falling, and measured accuracy collapses when a question is framed one way rather than another. A buyer in that position has no standard disclosure to rely on and no vendor obligation to produce one. The control has to sit on your side of the transaction, because there is nothing dependable on the other side to lean on.

What control would an accountant ask for?

Nothing exotic. The control an accountant would ask for is a review step that assumes the number is wrong until someone checks it, and a named person who signed. Applied to a language model, that looks like this.

  • Classify the task before you prompt. Decide whether the output has to be true or only has to be interesting, because those need different handling and the model will not tell you which one it just produced.
  • Verify consequential claims at the source. Open the case, the study, the filing, the invoice. A citation that formats correctly is evidence of formatting and nothing else.
  • Name the owner. Someone signs. The Nebraska court did not care whether AI was involved, and neither will any regulator, auditor or client who gets hurt.

Ambition is not the problem. Use these tools hard on the work where being approximately right quickly beats being exactly right slowly: drafts, options, first passes, rewrites, alternative framings, code you were going to test anyway. Keep the judgment and the review for the work where an error carries a price, and price the error honestly rather than assuming a confident tone means a checked one.

The industry's stated project is a model that does not hallucinate. The more useful project is a scoreboard that charges for confident error, which is what the researchers closest to the problem actually recommend.1 Until the benchmarks that rank these systems stop paying for a lucky guess, the systems will keep guessing, because guessing is what they are being paid to do. The review step is yours either way.