SyloSpace

Why do AI chatbots make things up?

AI chatbots make things up because hallucination is built into how large language models are trained and generate text, not a simple bug.

Updated 53 minutes ago5 min readVersion 2
CommentsFollow

Covers: This page explains why large language model chatbots generate false or fabricated information, covering the statistical and training-based causes of hallucination. It does not cover how to fix specific chatbot errors or compare individual products.

Also answers: Why do AI chatbots hallucinate? · Why does ChatGPT make up facts? · Why do AI chatbots lie? · What causes AI hallucinations?

A laptop displaying openai 4.0 chatbot interface
Photo: Brett Wharton

The short answer

Interpretation AI-prepared starting map

AI chatbots make things up because hallucination is a byproduct of how large language models are built and how they generate text, not a simple bug. Surveys define hallucination as content that is fluent and syntactically correct but factually inaccurate or unsupported by external evidence, or that diverges from the user's input or contradicts earlier context. Researchers trace the causes across the whole development lifecycle — data collection, architecture, training, and inference — and recent theory suggests hallucination is a built-in feature of how these systems generate text rather than a flaw that can simply be removed.123

What this rests on6 independent sources
  • Evidence 21
  • Opinion 1
  • Interpretation 1

Did this answer your question?

Be the first to vote

In brief

  1. Hallucination means fluent, syntactically correct output that is factually inaccurate or unsupported by external evidence — and it can also mean drifting from the user's input or contradicting earlier context.13

    Evidence-backed
  2. Causes span the whole development lifecycle: data, architecture, training, inference, and evaluation.12

    Evidence-backed
  3. Training data imbalance can cause 'knowledge overshadowing', where dominant conditions crowd out others; hallucination rate grows with the imbalance ratio and the length of the dominant condition description.4

    Evidence-backed
  4. Mitigation reduces hallucination but does not eliminate it, and retrieval augmentation has known limits.25

    Evidence-backed
  5. The most serious real-world harms are reported in legal and medical settings.2

    Evidence-backed

At a glance

The picture in numbers

Live · updated just now

Single paper on overshadowing-based mitigation

11.2–39.4%

0lowhigh

The high estimate is 3.5 times the low one.

Hallucination control achieved across different models and datasets4

The evidence behind it

6 sources
  • Reviews of many studies1
  • Other studies and data5

When it was published

Newest from 2026

20222026
Sources on this page by kind and year
SourceKindYear
Survey of Hallucination in Natural Language GenerationOther studies and data2022
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsOther studies and data2024
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language ModelsOther studies and data2023
Large Language Models Hallucination: A Comprehensive SurveyOther studies and data2025
AI Hallucinations in Large Language Model and Multimodal Models: A Systematic Review of Causes, Types, Detection, Mitigation, and Future DirectionsReviews of many studies2026
Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language ModelsOther studies and data2024

The community around it

No one has added to this page yet. Firsthand experience, a newer study or a different reading of the numbers would show up here, credited to you.

Shares and multiples are worked out from the figures the page states.

What it means for you

Which fits you?

Pick the situation closest to yours. Each answer says what it rests on.

If you want the short definition of hallucination

treat it as fluent, syntactically correct output that is factually inaccurate or unsupported by external evidence, and note that it also covers output that diverges from your input or contradicts earlier context.13

Evidence-backed

If you are trying to understand why a chatbot invents details

look at the whole pipeline rather than one stage: data collection, architecture, training, inference, and evaluation all contribute, and the mix differs by task.12

Evidence-backed

If you are asking a model about a niche topic alongside a popular one

be aware that knowledge overshadowing can occur: dominant conditions can crowd out the less common one, and the effect grows with the imbalance between them and with the length of the dominant condition's description.4

Evidence-backed

If you work in law or medicine

the surveys identify these as the settings where hallucination causes the most serious real-world harm, so treat outputs as requiring verification.2

Evidence-backed

If you are hoping a single fix will remove hallucination

expect reduction rather than elimination: mitigation techniques lessen the problem, and retrieval-augmented models have documented limitations in combating it.25

Evidence-backed

If you want to follow the research

the open directions flagged include hallucination in large vision-language models and understanding knowledge boundaries in LLM hallucinations.5

Evidence-backed

The full story · 5 chapters

01

What counts as making things up

AI summary:Hallucination is fluent, syntactically correct output that is factually wrong or unsupported, and can also mean drifting from the prompt or contradicting earlier context.

Evidence-backed

Evidence-backed: Across the surveys, hallucination is defined consistently in one respect: the output is fluent and syntactically correct but factually inaccurate or unsupported by external evidence. A separate survey adds two further forms — content that diverges from the user's input, and content that contradicts previously generated context or misaligns with established world knowledge. This matters because it means a chatbot can be wrong in ways that are not simply 'false facts': it can drift from what you asked, or contradict itself mid-answer.13

Evidence-backed

Evidence-backed: The problem is not new to chatbots. An earlier survey documented that deep-learning-based generation was already prone to producing unintended text in tasks such as abstractive summarization, dialogue generation, and data-to-text generation, degrading system performance and failing user expectations. What changed with large language models is scope: because they are open-ended and general-purpose, their hallucinations present challenges that differ from those of earlier task-specific models.65

02

Where it comes from

AI summary:Causes span the whole development lifecycle, and one theory argues hallucination is a built-in feature of how these systems generate text.

Evidence-backed

Evidence-backed: One survey analyses root causes across the entire development lifecycle — data collection, architecture design, and inference — and notes that hallucinations also emerge differently depending on the natural language generation task. Another survey reaches a similar conclusion from a different angle: hallucinations come from several factors working together, spanning the data, the system's design, how it is trained, how it generates answers, and how it is tested.12

Evidence-backed

Evidence-backed: A more specific mechanism is 'knowledge overshadowing'. When a model is queried with multiple conditions, some conditions overshadow others and the output is hallucinated. The researchers trace this partly to training data imbalance, verified across pretrained and fine-tuned models and across a wide range of model families and sizes, and interpret it theoretically as over-generalisation of the dominant conditions. They report that the hallucination rate grows with both the imbalance ratio between the popular and unpopular condition and the length of the dominant condition description, consistent with a derived generalisation bound.4

Evidence-backed

Evidence-backed: The 2026 review goes further and argues that recent theory treats hallucination as a built-in feature of how these systems generate text, not a flaw that can simply be removed. That is a stronger claim than the other surveys make, and it is the main point of disagreement in this material.2

03

How many kinds, and where it hurts most

AI summary:One review lists eight hallucination types, and the most serious real-world harms are reported in legal and medical settings.

Evidence-backed

Evidence-backed: The 2026 review identifies eight types of hallucination and reports that these types vary in how hard they are to detect and correct. It also states that the most serious real-world harms show up in legal and medical settings — domains where factual accuracy is the point.2

Evidence-backed

Evidence-backed: Other surveys organise the phenomenon differently, offering taxonomies of hallucination types and of evaluation benchmarks rather than a fixed count. One survey explicitly frames its taxonomy as new for the LLM era, and another pairs taxonomies of phenomena with taxonomies of benchmarks.53

04

Can it be stopped?

AI summary:Mitigation reduces hallucination but does not eliminate it, and retrieval augmentation has known limits.

Evidence-backed

Evidence-backed: The consistent answer across the surveys is: reduced, not eliminated. The 2026 review states plainly that techniques used to reduce hallucinations only lessen the problem — they do not get rid of it completely. Detection methods have improved, but that is a separate matter from prevention.2

Evidence-backed

Evidence-backed: One concrete mitigation approach uses the overshadowing conditions themselves as a signal to catch hallucination before it is produced, combined with a training-free self-contrastive decoding method to alleviate it during inference. The authors report up to 82% F1 for hallucination anticipation and 11.2% to 39.4% hallucination control across different models and datasets.4

Evidence-backed

Evidence-backed: Retrieval-augmented LLMs are a widely discussed route, but one survey specifically examines their current limitations in combating hallucinations and offers insights for building more robust information-retrieval systems — that is, retrieval helps but is not a complete answer.5

05

What readers think

AI summary:Readers can register how they approach the problem; this is a preference signal, not evidence about hallucination.

Participant opinion

Participant opinion: Readers of this page can register how they approach the problem. This is a preference signal, not evidence about how hallucination works.

Readers' pollNo answers yet

How often do you double-check factual claims made by AI chatbots?

How often do you double-check factual claims made by AI chatbots?

Your individual answer is private. Only totals are shown.

Your turn

Have your say

Quick votes, open to everyone. See where you stand the moment you vote. Only totals are ever shown.

How do you feel about this?

No votes yet

Quick questions from connected pages

Before you go

What to remember

The few things worth keeping from this page.

  1. Hallucination means fluent, syntactically correct output that is factually inaccurate or unsupported by external evidence — and it can also mean drifting from the user's input or contradicting earlier context.

  2. Causes span the whole development lifecycle: data, architecture, training, inference, and evaluation.

  3. Training data imbalance can cause 'knowledge overshadowing', where dominant conditions crowd out others; hallucination rate grows with the imbalance ratio and the length of the dominant condition description.

This answer keeps changing

When new evidence or a better source comes in, this page is updated (it's on version 2, last changed 53 minutes ago). Follow it to be told when that happens.

Up nextDo AI companion chatbots affect loneliness?Do AI companion chatbots reduce or worsen loneliness, and what does the evidence say about their effects on social connection?

Ask this Sylo

Still wondering about something?

Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.

Behind this page

Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.

Discussion

Nobody has added anything yet. If you have experience, evidence or a different view, you could be the first.

Sources

Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.

  1. 1
    Large Language Models Hallucination: A Comprehensive Survey
    arXiv (Cornell University) (Alansari & Luqman)Published Oct 5, 2025Checked Oct 10, 2026
    “Hallucination refers to the generation of content by an LLM that is fluent and syntactically correct but factually inaccurate or unsupported by external evidence. Hallucinations undermine the reliability and trustworthiness of LLMs, especially in domains requiring factual accuracy. This survey provides a comprehensive review of research on hallucination in LLMs, with a focus on causes, detection, and mitigation. We first present a taxonomy of hallucination types and analyze their root causes across the entire LLM development lifecycle, from data collection and architecture design to inference. We further examine how hallucinations emerge in key natural language generation tasks. Building on this foundation, we introduce a structured taxonomy of detection approaches and another taxonomy of mitigation strategies. We also analyze the strengths and limitations of current detection and mitigation approaches and review existing evaluation benchmarks and metrics used to quantify LLMs hallucinations. Finally, we outline key open challenges and promising directions for future research, providing a foundation for the development of more truthful and trustworthy LLMs.”
  2. 2
    AI Hallucinations in Large Language Model and Multimodal Models: A Systematic Review of Causes, Types, Detection, Mitigation, and Future Directions
    International Journal of Computer Sciences and Engineering (Singh & Banerjee)Published Jun 30, 2026Checked Oct 10, 2026
    “Language models are the problem and language models need to be fixed. This paper explores a problem that earlier research had left unexplored. It does this by reviewing existing studies and combining a mathematical theory with a clear step-, by-step approach. The researchers followed a process to find studies published between 2018 and 2026. They applied rules to screen these studies and narrowed them down to a smaller group. This group was then examined across eight themes. What they found is that AI “hallucinations” (when AI makes up false information) come from several factors working together: the data, the system's design, how it's trained, how it generates answers, and how it's tested. Recent theory suggests hallucination is a built-in feature of how these AI systems generate text, not a flaw that can simply be removed. They identified eight types of hallucinations, and these types vary in how hard they are to detect and correct. Some detection methods have improved, but the techniques used to reduce hallucinations only lessen the problem—they don't get rid of it completely. The most serious real-world harms show up in legal and medical settings.”
  3. 3
    Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
    arXiv (Cornell University) (Zhang et al.)Published Sep 3, 2023Checked Oct 10, 2026
    “While large language models (LLMs) have demonstrated remarkable capabilities across a range of downstream tasks, a significant concern revolves around their propensity to exhibit hallucinations: LLMs occasionally generate content that diverges from the user input, contradicts previously generated context, or misaligns with established world knowledge. This phenomenon poses a substantial challenge to the reliability of LLMs in real-world scenarios. In this paper, we survey recent efforts on the detection, explanation, and mitigation of hallucination, with an emphasis on the unique challenges posed by LLMs. We present taxonomies of the LLM hallucination phenomena and evaluation benchmarks, analyze existing approaches aiming at mitigating LLM hallucination, and discuss potential directions for future research.”
  4. 4
    Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models
    arXiv (Cornell University) (Zhang et al.)Published Jul 10, 2024Checked Oct 10, 2026
    “We coin this phenomenon as ``knowledge overshadowing'': when we query knowledge from a language model with multiple conditions, some conditions overshadow others, leading to hallucinated outputs. This phenomenon partially stems from training data imbalance, which we verify on both pretrained models and fine-tuned models, over a wide range of LM model families and sizes.From a theoretical point of view, knowledge overshadowing can be interpreted as over-generalization of the dominant conditions (patterns). We show that the hallucination rate grows with both the imbalance ratio (between the popular and unpopular condition) and the length of dominant condition description, consistent with our derived generalization bound. Finally, we propose to utilize overshadowing conditions as a signal to catch hallucination before it is produced, along with a training-free self-contrastive decoding method to alleviate hallucination during inference. Our proposed approach showcases up to 82% F1 for hallucination anticipation and 11.2% to 39.4% hallucination control, with different models and datasets.”
  5. 5
    A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
    ACM Transactions on Information Systems (Huang et al.)Published Nov 20, 2024Checked Oct 10, 2026
    “Given the open-ended general-purpose attributes inherent to LLMs, LLM hallucinations present distinct challenges that diverge from prior task-specific models. This divergence highlights the urgency for a nuanced understanding and comprehensive overview of recent advances in LLM hallucinations. In this survey, we begin with an innovative taxonomy of hallucination in the era of LLM and then delve into the factors contributing to hallucinations. Subsequently, we present a thorough overview of hallucination detection methods and benchmarks. Our discussion then transfers to representative methodologies for mitigating LLM hallucinations. Additionally, we delve into the current limitations faced by retrieval-augmented LLMs in combating hallucinations, offering insights for developing more robust IR systems. Finally, we highlight the promising research directions on LLM hallucinations, including hallucination in large vision-language models and understanding of knowledge boundaries in LLM hallucinations.”
  6. 6
    Survey of Hallucination in Natural Language Generation
    ACM Computing Surveys (Ji et al.)Published Nov 17, 2022Checked Oct 10, 2026
    “This advancement has led to more fluent and coherent NLG, leading to improved development in downstream tasks such as abstractive summarization, dialogue generation, and data-to-text generation. However, it is also apparent that deep learning based generation is prone to hallucinate unintended text, which degrades the system performance and fails to meet user expectations in many real-world scenarios. To address this issue, many studies have been presented in measuring and mitigating hallucinated texts, but these have never been reviewed in a comprehensive manner before. In this survey, we thus provide a broad overview of the research progress and challenges in the hallucination problem in NLG. The survey is organized into two parts: (1) a general overview of metrics, mitigation methods, and future directions, and (2) an overview of task-specific research progress on hallucinations in the following downstream tasks, namely abstractive summarization, dialogue generation, generative question answering, data-to-text generation, and machine translation. This survey serves to facilitate collaborative efforts among researchers in tackling the challenge of hallucinated texts in NLG.”

How it changed

Published 1 time since Oct 10, 2026.

  1. Version 2Oct 10, 2026Live now

    AI-prepared Starting Map from live research.

    • First published version.
Every version, side by side

Help improve it

The brief is open about what's uncertain. These are the specific gaps that new material would fill.

  • “What readers think” has no evidence or firsthand experience yet

    It's a synthesis for now. Evidence or experience would show whether it holds.

Open questions

  • Is hallucination an inherent property of how language models generate text, as the 2026 review argues, or a defect that better data, training, and decoding can largely remove?

    No answers yet

  • How far does knowledge overshadowing generalise beyond the models and datasets tested, and does it explain hallucinations that are not tied to competing conditions?

    No answers yet

  • What specifically limits retrieval-augmented models in reducing hallucination, and what would a more robust retrieval system need to do?

    No answers yet

  • Why do the most serious harms concentrate in legal and medical settings, and do other high-stakes domains show similar patterns?

    No answers yet

Around this topic

Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.

Ask this Sylo

Answers only from “Why do AI chatbots make things up?”

Ask anything about this page. The AI reads only its reviewed brief, sources and contributions, cites what it used, and says when the page doesn't cover something.