Why do AI chatbots make things up?
AI chatbots make things up because hallucination is built into how large language models are trained and generate text, not a simple bug.
Covers: This page explains why large language model chatbots generate false or fabricated information, covering the statistical and training-based causes of hallucination. It does not cover how to fix specific chatbot errors or compare individual products.
Also answers: Why do AI chatbots hallucinate? · Why does ChatGPT make up facts? · Why do AI chatbots lie? · What causes AI hallucinations?
- One page for this question6 other ways of asking lead here
- 6 independent sourcesEvery claim links to what supports it
- Joins the mapLinked as related pages appear
- Clean discussionScreened before anything appears
The short answer
Interpretation AI-prepared starting mapAI chatbots make things up because hallucination is a byproduct of how large language models are built and how they generate text, not a simple bug. Surveys define hallucination as content that is fluent and syntactically correct but factually inaccurate or unsupported by external evidence, or that diverges from the user's input or contradicts earlier context. Researchers trace the causes across the whole development lifecycle — data collection, architecture, training, and inference — and recent theory suggests hallucination is a built-in feature of how these systems generate text rather than a flaw that can simply be removed.123
- Evidence 21
- Opinion 1
- Interpretation 1
Did this answer your question?
Be the first to voteIn brief
Training data imbalance can cause 'knowledge overshadowing', where dominant conditions crowd out others; hallucination rate grows with the imbalance ratio and the length of the dominant condition description.4
Evidence-backed
At a glance
The picture in numbers
Live · updated just now
11.2–39.4%
The high estimate is 3.5 times the low one.
The evidence behind it
6 sources- Reviews of many studies1
- Other studies and data5
When it was published
Newest from 2026
| Source | Kind | Year |
|---|---|---|
| Survey of Hallucination in Natural Language Generation | Other studies and data | 2022 |
| A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions | Other studies and data | 2024 |
| Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models | Other studies and data | 2023 |
| Large Language Models Hallucination: A Comprehensive Survey | Other studies and data | 2025 |
| AI Hallucinations in Large Language Model and Multimodal Models: A Systematic Review of Causes, Types, Detection, Mitigation, and Future Directions | Reviews of many studies | 2026 |
| Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models | Other studies and data | 2024 |
The community around it
No one has added to this page yet. Firsthand experience, a newer study or a different reading of the numbers would show up here, credited to you.
Shares and multiples are worked out from the figures the page states.
What it means for you
Which fits you?
Pick the situation closest to yours. Each answer says what it rests on.
If you want the short definition of hallucination
treat it as fluent, syntactically correct output that is factually inaccurate or unsupported by external evidence, and note that it also covers output that diverges from your input or contradicts earlier context.13
Evidence-backedIf you are trying to understand why a chatbot invents details
look at the whole pipeline rather than one stage: data collection, architecture, training, inference, and evaluation all contribute, and the mix differs by task.12
Evidence-backedIf you are asking a model about a niche topic alongside a popular one
be aware that knowledge overshadowing can occur: dominant conditions can crowd out the less common one, and the effect grows with the imbalance between them and with the length of the dominant condition's description.4
Evidence-backedIf you work in law or medicine
the surveys identify these as the settings where hallucination causes the most serious real-world harm, so treat outputs as requiring verification.2
Evidence-backedIf you are hoping a single fix will remove hallucination
expect reduction rather than elimination: mitigation techniques lessen the problem, and retrieval-augmented models have documented limitations in combating it.25
Evidence-backedIf you want to follow the research
the open directions flagged include hallucination in large vision-language models and understanding knowledge boundaries in LLM hallucinations.5
Evidence-backedThe full story · 5 chapters
01
What counts as making things up
AI summary:Hallucination is fluent, syntactically correct output that is factually wrong or unsupported, and can also mean drifting from the prompt or contradicting earlier context.
Evidence-backed: Across the surveys, hallucination is defined consistently in one respect: the output is fluent and syntactically correct but factually inaccurate or unsupported by external evidence. A separate survey adds two further forms — content that diverges from the user's input, and content that contradicts previously generated context or misaligns with established world knowledge. This matters because it means a chatbot can be wrong in ways that are not simply 'false facts': it can drift from what you asked, or contradict itself mid-answer.13
Evidence-backed: The problem is not new to chatbots. An earlier survey documented that deep-learning-based generation was already prone to producing unintended text in tasks such as abstractive summarization, dialogue generation, and data-to-text generation, degrading system performance and failing user expectations. What changed with large language models is scope: because they are open-ended and general-purpose, their hallucinations present challenges that differ from those of earlier task-specific models.65
02
Where it comes from
AI summary:Causes span the whole development lifecycle, and one theory argues hallucination is a built-in feature of how these systems generate text.
Evidence-backed: One survey analyses root causes across the entire development lifecycle — data collection, architecture design, and inference — and notes that hallucinations also emerge differently depending on the natural language generation task. Another survey reaches a similar conclusion from a different angle: hallucinations come from several factors working together, spanning the data, the system's design, how it is trained, how it generates answers, and how it is tested.12
Evidence-backed: A more specific mechanism is 'knowledge overshadowing'. When a model is queried with multiple conditions, some conditions overshadow others and the output is hallucinated. The researchers trace this partly to training data imbalance, verified across pretrained and fine-tuned models and across a wide range of model families and sizes, and interpret it theoretically as over-generalisation of the dominant conditions. They report that the hallucination rate grows with both the imbalance ratio between the popular and unpopular condition and the length of the dominant condition description, consistent with a derived generalisation bound.4
Evidence-backed: The 2026 review goes further and argues that recent theory treats hallucination as a built-in feature of how these systems generate text, not a flaw that can simply be removed. That is a stronger claim than the other surveys make, and it is the main point of disagreement in this material.2
03
How many kinds, and where it hurts most
AI summary:One review lists eight hallucination types, and the most serious real-world harms are reported in legal and medical settings.
Evidence-backed: The 2026 review identifies eight types of hallucination and reports that these types vary in how hard they are to detect and correct. It also states that the most serious real-world harms show up in legal and medical settings — domains where factual accuracy is the point.2
Evidence-backed: Other surveys organise the phenomenon differently, offering taxonomies of hallucination types and of evaluation benchmarks rather than a fixed count. One survey explicitly frames its taxonomy as new for the LLM era, and another pairs taxonomies of phenomena with taxonomies of benchmarks.53
04
Can it be stopped?
AI summary:Mitigation reduces hallucination but does not eliminate it, and retrieval augmentation has known limits.
Evidence-backed: The consistent answer across the surveys is: reduced, not eliminated. The 2026 review states plainly that techniques used to reduce hallucinations only lessen the problem — they do not get rid of it completely. Detection methods have improved, but that is a separate matter from prevention.2
Evidence-backed: One concrete mitigation approach uses the overshadowing conditions themselves as a signal to catch hallucination before it is produced, combined with a training-free self-contrastive decoding method to alleviate it during inference. The authors report up to 82% F1 for hallucination anticipation and 11.2% to 39.4% hallucination control across different models and datasets.4
Evidence-backed: Retrieval-augmented LLMs are a widely discussed route, but one survey specifically examines their current limitations in combating hallucinations and offers insights for building more robust information-retrieval systems — that is, retrieval helps but is not a complete answer.5
05
What readers think
AI summary:Readers can register how they approach the problem; this is a preference signal, not evidence about hallucination.
Participant opinion: Readers of this page can register how they approach the problem. This is a preference signal, not evidence about how hallucination works.
How often do you double-check factual claims made by AI chatbots?
Your individual answer is private. Only totals are shown.
Your turn
Have your say
Quick votes, open to everyone. See where you stand the moment you vote. Only totals are ever shown.
How do you feel about this?
No votes yetQuick questions from connected pages
Before you go
What to remember
The few things worth keeping from this page.
Hallucination means fluent, syntactically correct output that is factually inaccurate or unsupported by external evidence — and it can also mean drifting from the user's input or contradicting earlier context.
Causes span the whole development lifecycle: data, architecture, training, inference, and evaluation.
Training data imbalance can cause 'knowledge overshadowing', where dominant conditions crowd out others; hallucination rate grows with the imbalance ratio and the length of the dominant condition description.
Your reading
0 of 5 chaptersThis answer keeps changing
When new evidence or a better source comes in, this page is updated (it's on version 2, last changed 53 minutes ago). Follow it to be told when that happens.
Ask this Sylo
Still wondering about something?
Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.
More on AI
Everything on AI ›How much energy does ChatGPT use?
How much energy does ChatGPT use per query, and how does that compare with other everyday activities?
Does AI help students learn or make them lazier?
Does using AI tools help students learn more effectively, or does it make them lazier and undermine their learning?
Is using ChatGPT for homework cheating?
How accurate is ChatGPT?
How accurate is ChatGPT, and what does research show about its error rates and reliability?
Is AI dangerous?
Is artificial intelligence dangerous, and what does the evidence say about its risks?
When will we have AGI?
When will we have artificial general intelligence (AGI)?
Behind this page
Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.
Discussion
Sources
Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.
- 1Large Language Models Hallucination: A Comprehensive SurveyarXiv (Cornell University) (Alansari & Luqman)Published Oct 5, 2025Checked Oct 10, 2026
“Hallucination refers to the generation of content by an LLM that is fluent and syntactically correct but factually inaccurate or unsupported by external evidence. Hallucinations undermine the reliability and trustworthiness of LLMs, especially in domains requiring factual accuracy. This survey provides a comprehensive review of research on hallucination in LLMs, with a focus on causes, detection, and mitigation. We first present a taxonomy of hallucination types and analyze their root causes across the entire LLM development lifecycle, from data collection and architecture design to inference. We further examine how hallucinations emerge in key natural language generation tasks. Building on this foundation, we introduce a structured taxonomy of detection approaches and another taxonomy of mitigation strategies. We also analyze the strengths and limitations of current detection and mitigation approaches and review existing evaluation benchmarks and metrics used to quantify LLMs hallucinations. Finally, we outline key open challenges and promising directions for future research, providing a foundation for the development of more truthful and trustworthy LLMs.”
- 2AI Hallucinations in Large Language Model and Multimodal Models: A Systematic Review of Causes, Types, Detection, Mitigation, and Future DirectionsInternational Journal of Computer Sciences and Engineering (Singh & Banerjee)Published Jun 30, 2026Checked Oct 10, 2026
“Language models are the problem and language models need to be fixed. This paper explores a problem that earlier research had left unexplored. It does this by reviewing existing studies and combining a mathematical theory with a clear step-, by-step approach. The researchers followed a process to find studies published between 2018 and 2026. They applied rules to screen these studies and narrowed them down to a smaller group. This group was then examined across eight themes. What they found is that AI “hallucinations” (when AI makes up false information) come from several factors working together: the data, the system's design, how it's trained, how it generates answers, and how it's tested. Recent theory suggests hallucination is a built-in feature of how these AI systems generate text, not a flaw that can simply be removed. They identified eight types of hallucinations, and these types vary in how hard they are to detect and correct. Some detection methods have improved, but the techniques used to reduce hallucinations only lessen the problem—they don't get rid of it completely. The most serious real-world harms show up in legal and medical settings.”
- 3Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language ModelsarXiv (Cornell University) (Zhang et al.)Published Sep 3, 2023Checked Oct 10, 2026
“While large language models (LLMs) have demonstrated remarkable capabilities across a range of downstream tasks, a significant concern revolves around their propensity to exhibit hallucinations: LLMs occasionally generate content that diverges from the user input, contradicts previously generated context, or misaligns with established world knowledge. This phenomenon poses a substantial challenge to the reliability of LLMs in real-world scenarios. In this paper, we survey recent efforts on the detection, explanation, and mitigation of hallucination, with an emphasis on the unique challenges posed by LLMs. We present taxonomies of the LLM hallucination phenomena and evaluation benchmarks, analyze existing approaches aiming at mitigating LLM hallucination, and discuss potential directions for future research.”
- 4Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language ModelsarXiv (Cornell University) (Zhang et al.)Published Jul 10, 2024Checked Oct 10, 2026
“We coin this phenomenon as ``knowledge overshadowing'': when we query knowledge from a language model with multiple conditions, some conditions overshadow others, leading to hallucinated outputs. This phenomenon partially stems from training data imbalance, which we verify on both pretrained models and fine-tuned models, over a wide range of LM model families and sizes.From a theoretical point of view, knowledge overshadowing can be interpreted as over-generalization of the dominant conditions (patterns). We show that the hallucination rate grows with both the imbalance ratio (between the popular and unpopular condition) and the length of dominant condition description, consistent with our derived generalization bound. Finally, we propose to utilize overshadowing conditions as a signal to catch hallucination before it is produced, along with a training-free self-contrastive decoding method to alleviate hallucination during inference. Our proposed approach showcases up to 82% F1 for hallucination anticipation and 11.2% to 39.4% hallucination control, with different models and datasets.”
- 5A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsACM Transactions on Information Systems (Huang et al.)Published Nov 20, 2024Checked Oct 10, 2026
“Given the open-ended general-purpose attributes inherent to LLMs, LLM hallucinations present distinct challenges that diverge from prior task-specific models. This divergence highlights the urgency for a nuanced understanding and comprehensive overview of recent advances in LLM hallucinations. In this survey, we begin with an innovative taxonomy of hallucination in the era of LLM and then delve into the factors contributing to hallucinations. Subsequently, we present a thorough overview of hallucination detection methods and benchmarks. Our discussion then transfers to representative methodologies for mitigating LLM hallucinations. Additionally, we delve into the current limitations faced by retrieval-augmented LLMs in combating hallucinations, offering insights for developing more robust IR systems. Finally, we highlight the promising research directions on LLM hallucinations, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations.”
- 6Survey of Hallucination in Natural Language GenerationACM Computing Surveys (Ji et al.)Published Nov 17, 2022Checked Oct 10, 2026
“This advancement has led to more fluent and coherent NLG, leading to improved development in downstream tasks such as abstractive summarization, dialogue generation, and data-to-text generation. However, it is also apparent that deep learning based generation is prone to hallucinate unintended text, which degrades the system performance and fails to meet user expectations in many real-world scenarios. To address this issue, many studies have been presented in measuring and mitigating hallucinated texts, but these have never been reviewed in a comprehensive manner before. In this survey, we thus provide a broad overview of the research progress and challenges in the hallucination problem in NLG. The survey is organized into two parts: (1) a general overview of metrics, mitigation methods, and future directions, and (2) an overview of task-specific research progress on hallucinations in the following downstream tasks, namely abstractive summarization, dialogue generation, generative question answering, data-to-text generation, and machine translation. This survey serves to facilitate collaborative efforts among researchers in tackling the challenge of hallucinated texts in NLG.”
How it changed
Published 1 time since Oct 10, 2026.
- Version 2Oct 10, 2026Live now
AI-prepared Starting Map from live research.
- First published version.
Help improve it
The brief is open about what's uncertain. These are the specific gaps that new material would fill.
“What readers think” has no evidence or firsthand experience yet
It's a synthesis for now. Evidence or experience would show whether it holds.
Open questions
Is hallucination an inherent property of how language models generate text, as the 2026 review argues, or a defect that better data, training, and decoding can largely remove?
No answers yet
How far does knowledge overshadowing generalise beyond the models and datasets tested, and does it explain hallucinations that are not tied to competing conditions?
No answers yet
What specifically limits retrieval-augmented models in reducing hallucination, and what would a more robust retrieval system need to do?
No answers yet
Why do the most serious harms concentrate in legal and medical settings, and do other high-stakes domains show similar patterns?
No answers yet
Around this topic
Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.