Call · 15 min
GlossaryMattia Esposito26 September 20267-minute read

AI hallucinations. When a model makes something up, and says it with the same confidence as the truth.

A hallucination is a false statement produced by an artificial intelligence model with the same fluency and confidence as a true one: a figure, a quotation, a rule, a price that doesn't exist. The danger lies in the confidence with which it's stated.

In brief

OpenAI defines them as plausible but false statements, and in September 2025 it traced them to training and evaluation procedures that reward guessing over admitting uncertainty.

A model that knows how to abstain gets far less wrong: in the table OpenAI published, a model that abstains on 52% of questions gets 26% wrong, while one that abstains on 1% gets 75% wrong.

Responsibility stays with the business. In 2024 a Canadian tribunal ruled against an airline that didn't want to answer for the wrong information its chatbot had given.

This entry is part of the AI and automation glossary, where the term appears as hallucination. Here the definition goes further: what they are, why a model makes things up, how often it happens according to the published measurements, who's responsible, and what makes them rare.

What hallucinations are

In an article of 5 September 2025, OpenAI defines hallucinations as “plausible but false statements generated by language models”. They can be about anything: a date, a name, a rule, the title of a book, a number in a document the model was only meant to read.

The example the authors themselves use is revealing. Asked for the title of one of their doctoral theses, a very popular chatbot confidently gave three different answers, all wrong; asked for that researcher's date of birth, it gave three different dates, all wrong too.

Why a model makes things up

An LLM learns to predict the next word from huge amounts of text, and nothing in that text labels what's true and what's false. Patterns, such as spelling, are learned well. Rare, arbitrary facts, like the birthday of someone little known, can't be inferred from any pattern, and that's where the model guesses.

The second cause lies in evaluations. According to OpenAI, most tests measure only right answers, and saying “I don't know” scores zero: as in a multiple-choice quiz, guessing pays. A model trained to do well on those tests learns to always answer, even when it doesn't know.

How often they get it wrong: the published measurements

The published measurements vary a lot with the task and with whether the model can abstain, and none of them predicts your case. Taken together, they show where the risk lies: in the questions the model can't answer but answers anyway.

MeasureWhat it foundOn which task
OpenAISeptember 2025

A model that abstains on 52% of questions gets 26% wrong; one that abstains on 1% gets 75% wrong.

Questions with a single right answer, from the SimpleQA test.

Stanford RegLab2024

Three legal research tools marketed as hallucination-free get it wrong between 17% and 33% of the time.

Questions on US law, with documents available to the system.

Itriatest bench, 4 September 2026

An unreadable PDF, deliberately placed in a batch of 30 documents, set aside as “other” without a single invented field.

Extracting data from orders, delivery notes and invoices.

The first row holds the most useful lesson. The two OpenAI models give almost the same share of right answers, 22% against 24%, but the first gets 26 in a hundred wrong instead of 75, because it abstains when it doesn't know. In a business, an “I don't know” costs a question passed to a colleague; an invented answer can cost a customer.

Who's responsible: a case decided by a tribunal

On 14 February 2024 the Civil Resolution Tribunal of British Columbia, in Canada, decided Moffatt v. Air Canada. The chatbot on the airline's website had given a customer wrong information about a bereavement fare, and the airline argued it couldn't be held responsible for what the chatbot said.

The tribunal rejected the argument, in paragraph 27: “It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot”. The airline was ordered to refund the difference in fare. As far as the customer is concerned, what your assistant says, you say.

Can they be avoided?

They can be made rare, and the key is abstaining. OpenAI puts it this way: accuracy will never reach 100%, because some real-world questions have no answer that can be worked out, but hallucinations aren't inevitable, because a model can abstain when it's unsure. A business system needs three safeguards.

The right sources at the moment of the question. A system that answers from the business's own documents, using the technique called RAG, can say which document each answer comes from, and the person reading it can check.

Permission to say “I don't know”. The system has to be built so that, when a source is missing, it stops and passes the question to a person, instead of finishing the answer anyway.

A person wherever a mistake is costly. Prices, terms, deadlines and anything that commits the business go through a human in the loop before sending.

How Itria handles them

In Itria's systems the rule is written before the code: when there's no evidence, the system stops and alerts a person. On the test bench of 4 September 2026, the unreadable document placed there deliberately was set aside as “other”, without a single invented field, and across the whole batch 234 of 237 extracted fields were correct.

The knowledge base chatbot answers from documents the business has approved and says it's a system. The principles are set out in Ethics, and the third matters most here: the machine prepares, a person always decides.

Related terms

LLM

The language model that produces the text, and with the text the hallucinations. Understanding how it works explains why it makes things up.

Human in the loop

The person who approves before sending. It's the last line of defence, and it belongs wherever a mistake is costly.

RAG

The technique that gives the model the right documents at the moment of the question, so it answers from them.

Knowledge base

The set of approved documents a system can answer from. If it's out of date, the answers are wrong even without hallucinations.

Questions and answers

What are AI hallucinations?

They're false statements produced by an artificial intelligence model with the same fluency and confidence as true ones: a figure, a quotation, a rule, a name or a number that doesn't exist.

OpenAI defines them as plausible but false statements generated by language models. The danger lies in the confidence with which they're stated, which makes them hard to spot.

Why does artificial intelligence make up answers?

For two reasons, according to research published by OpenAI on 5 September 2025. First, a language model learns to predict the next word, and rare, arbitrary facts can't be inferred from any pattern.

Second, most tests reward only right answers, and saying I don't know scores zero, so models learn to guess rather than admit uncertainty.

Can AI hallucinations be avoided?

They can be made rare. OpenAI writes that accuracy will never reach 100%, but that hallucinations aren't inevitable, because a model can abstain when it's unsure.

A business system needs three safeguards: give the model the right sources at the moment of the question, build it to stop when a source is missing, and put a person before sending wherever a mistake is costly.

Who is responsible if a chatbot gives wrong information?

In Moffatt v. Air Canada, decided on 14 February 2024 by the Civil Resolution Tribunal of British Columbia, the airline argued it wasn't responsible for what its chatbot said. The tribunal ruled the opposite: the business is responsible for all the information on its website, whether it comes from a page or a chatbot.

It's a Canadian decision, cited for the principle; for a specific case, ask the people who advise your business.

What are some examples of AI hallucinations?

A citation of a law or ruling that doesn't exist, an invented statistic, the wrong title for a book or thesis, a false date of birth, a rate or commercial term the business never offered, a total filled in from an unreadable document.

OpenAI describes how a popular chatbot gave three different titles, all wrong, for the doctoral thesis of one of its researchers.

Notes on sources

  1. The definition, the thesis example, the two causes and the table with 52%, 26%, 1% and 75% come from OpenAI, Why language models hallucinate, 5 September 2025, read on 26 September 2026. It's the view of the people who build the models, published alongside a research paper; the table comes from the system card of one of their models.
  2. The 17% and 33% come from the pre-registered evaluation by Stanford's RegLab group, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, 2024: three US legal tools on questions of law, a harder task than an SME's.
  3. The decision is Moffatt v. Air Canada, 2024 BCCRT 149, of 14 February 2024, read on 26 September 2026 on the tribunal's website; the quotation is from paragraph 27. It's a decision by a Canadian small-claims tribunal: it isn't a precedent in Italy, and we cite it for the principle.
  4. The unreadable document and the 234 correct fields out of 237 come from an Itria measurement on a test bench on 4 September 2026, 30 documents. It's a test system, not work carried out for a client.
·The next step

A reliable system knows when to stop. It's the first thing to ask anyone pitching it to you.

The first step with Itria is a fifteen-minute video call: we look at where a system could work for you, and at which points it has to stop and hand over to a person. Drop us a line about what's slowing you down. We'll make the first move: we'll look at what a customer sees when they search for you, and tell you what we found. Even if we never end up working together.