How accurate is ChatGPT?
How accurate ChatGPT is depends heavily on the task, the question type, and which model version you use, and it still makes things up.
Covers: Published evaluations of ChatGPT's factual accuracy, hallucination rates and error patterns across tasks and model versions, plus what affects reliability. Does not cover how to use ChatGPT for specific tasks or compare it with other chatbots in depth.
Also answers: How reliable is ChatGPT? · Does ChatGPT make mistakes? · How often is ChatGPT wrong?
- One page for this question6 other ways of asking lead here
- 7 independent sourcesEvery claim links to what supports it
- Joins the mapLinked as related pages appear
- Clean discussionScreened before anything appears
Before you read, make a guess
Fill in the blank: ?% of ChatGPT responses in HaluEval contained fabricated or unverifiable content
Drag the slider to fill in the blank
The short answer
Evidence-backed AI-prepared starting mapPublished evaluations show ChatGPT's factual accuracy varies widely by task, question type and model version. On medical licensing-style multiple-choice exams, reported accuracy ranges from 42.7% (GPT-3.5, ophthalmology) and 42.8% (GPT-3.5, Japanese NMLE) up to 81.5% (GPT-4, Japanese NMLE) and 86–94% for current-generation models on European anesthesiology exams. On patient-facing material, 88% of ChatGPT responses to 50 total knee replacement FAQs were rated accurate (mean 4.6/5) and 100% relevant (4.9/5). On clinical pharmacy factual questions, ChatGPT answered 79% accurately versus 66% for pharmacists, with 95% concordance and high reproducibility (>92%). Hallucination is a persistent, measured problem: about 19.5% of ChatGPT responses in the HaluEval benchmark contained fabricated or unverifiable content, and hallucination rates of 11–20.1% were found across current LLMs on anesthesiology exams.123456
- Evidence 21
- Interpretation 1
Did this answer your question?
Be the first to voteIn brief
The evidence base is mostly benchmark and simulated-workflow studies (89%), so real-world reliability under supervision remains less well characterized.7
Evidence-backed
At a glance
The picture in numbers
Live · updated just now
- GPT-3.5, ophthalmology42.7%
- GPT-3.5, Japanese NMLE42.8%
- GPT-4, Japanese NMLE81.5%
- Current models, anesthesiology86–94%
- ChatGPT79%
- Pharmacists66%
19.5%
20 in every 100
11–20.1%
The high estimate is 1.8 times the low one.
The evidence behind it
7 sources- Reviews of many studies1
- Other studies and data6
When it was published
Newest from 2026
| Source | Kind | Year |
|---|---|---|
| Evaluating the Performance of ChatGPT in Ophthalmology | Other studies and data | 2023 |
| Evaluating the accuracy and relevance of ChatGPT responses to frequently asked questions regarding total knee replacement | Other studies and data | 2024 |
| Accuracy of ChatGPT on Medical Questions in the National Medical Licensing Examination in Japan: Evaluation Study | Other studies and data | 2023 |
| Performance of ChatGPT on Factual Knowledge Questions Regarding Clinical Pharmacy | Other studies and data | 2024 |
| Performance and Hallucination Analysis of Large Language Models on European Anesthesiology Examinations: Cross-Sectional Comparative Study. | Other studies and data | 2026 |
| HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models | Other studies and data | 2023 |
| Large language models in healthcare: applications, evaluation frameworks, and governance pathways - a scoping review and multidimensional framework. | Reviews of many studies | 2026 |
The community around it
No one has added to this page yet. Firsthand experience, a newer study or a different reading of the numbers would show up here, credited to you.
Shares and multiples are worked out from the figures the page states.
What it means for you
Which fits you?
Pick the situation closest to yours. Each answer says what it rests on.
If you are using ChatGPT for patient education material
treat responses as a starting point that needs verification: 88% of knee replacement FAQ answers were rated accurate, but the authors warn it can occasionally provide inaccurate medical information and recommend supervision.3
Evidence-backedIf you are relying on ChatGPT for clinical or pharmacy facts
expect roughly 79% accuracy on clinical pharmacy factual questions with high reproducibility (>92%), but verify substantiation, since only 73% of responses had good or excellent substantiation.4
Evidence-backedIf you are using ChatGPT as a study or exam-preparation aid
current models exceed medical exam passing thresholds (86–94% on anesthesiology exams; GPT-4 at 81.5% on the Japanese licensing exam), but hallucination rates of 11–20.1% mean it should be a supervised educational tool, not an autonomous learning resource.62
Evidence-backedIf you are asking about a specialized medical subspecialty
expect lower accuracy: ophthalmology performance was worst in neuro-ophthalmology and ophthalmic pathology and intraocular tumors, and the authors suggest domain-specific pre-training may be needed.1
Evidence-backedIf you are deploying an LLM in a healthcare workflow
use structured oversight and systematic verification: the scoping review supports cautious deployment in selected tasks and flags hallucinated content, omitted critical information, bias, privacy and automation bias as risks.7
Evidence-backedIf you are comparing older reports of ChatGPT accuracy with current experience
check the model version: GPT-3.5 scored 42.8% on the Japanese licensing exam where GPT-4 scored 81.5%, and later models reached 86–94% on anesthesiology exams.26
Evidence-backedIf you need ChatGPT to recognize or avoid hallucinations in text
providing external knowledge or adding reasoning steps helped LLMs detect hallucinations in the HaluEval experiments.5
Evidence-backedThe full story · 4 chapters
01
Accuracy on medical and licensing-style exams
AI summary:On medical licensing-style exams, scores ranged from 42.7% for GPT-3.5 to 81.5% for GPT-4, with newer models reaching 86–94%.
Evidence-backed: On two 260-question simulated ophthalmology (OKAP) exams of easy-to-moderate difficulty, ChatGPT scored 55.8% and 42.7%. Performance varied by subspecialty, best in general medicine and worst in neuro-ophthalmology and ophthalmic pathology and intraocular tumors; the authors concluded that domain-specific pre-training may be needed for subspecialty performance.1
Evidence-backed: On the Japanese National Medical Licensing Examination, GPT-4 answered 81.5% (237/292) correctly versus 42.8% (125/292) for GPT-3.5 — a significant difference — and GPT-4 surpassed the >72% passing standard. The analysis excluded questions containing charts, which ChatGPT did not support, so this result applies to written questions only.2
Evidence-backed: On European anesthesiology examinations, current-generation LLMs achieved average success rates of 86% (SD 18%) to 94% (SD 10%), all exceeding the EDAIC part I passing threshold and representing substantial improvement over previously reported GPT-3.5 performance. Intermodel differences were significant overall (Friedman χ²₃=13.9; P=.003; W=0.02), with Gemini outperforming GPT-5 as the only pairwise difference.6
02
Accuracy on patient questions and pharmacy knowledge
AI summary:ChatGPT scored well on knee replacement FAQs and clinical pharmacy facts, though the authors still advise supervision and verification.
Evidence-backed: For 50 frequently asked questions about total knee replacement, 44/50 (88%) of ChatGPT responses were classified as accurate (mean Likert 4.6/5) and 50/50 (100%) as relevant (4.9/5). The authors note it is not infallible and can occasionally provide inaccurate medical information, recommending supervision and verification.3
Evidence-backed: On factual knowledge questions in clinical pharmacy, ChatGPT answered 79% accurately, compared with 66% for pharmacists in 2022. Concordance (no contradictions) was 95%, substantiation quality was good or excellent for 73% of questions, and reproducibility was consistently high both within and between days and across users (>92%).4
03
Hallucination rates and error patterns
AI summary:Hallucination is measurable and persistent, so the evidence supports supervised use rather than autonomous clinical reliance.
Evidence-backed: The HaluEval benchmark found ChatGPT likely to generate hallucinated content on specific topics by fabricating unverifiable information in about 19.5% of responses. Existing LLMs also struggled to recognize hallucinations in text, but providing external knowledge or adding reasoning steps helped them detect hallucinations.5
Evidence-backed: On European anesthesiology examinations, hallucination rates ranged from 11% (11/100) to 20.1% (44/219) with no significant differences between models. Despite high exam scores, the authors concluded that current LLMs continue to produce clinically relevant hallucinations, supporting their role as supervised educational tools rather than autonomous learning resources.6
Evidence-backed: A scoping review of LLMs in healthcare lists hallucinated content, omission of clinically critical information, demographic bias, privacy vulnerabilities, limited explainability and automation bias among safety-relevant risks, and notes that reported benefits concentrated on documentation efficiency, text quality and knowledge synthesis. It concludes that current evidence supports cautious deployment in selected healthcare tasks under structured oversight.7
04
What affects reliability
AI summary:Model version and task type drive accuracy, while repeated factual answers stayed stable and hallucination remained measurable.
Interpretation: Model version is a major factor: GPT-4 substantially outperformed GPT-3.5 on the same Japanese licensing questions (81.5% vs 42.8%), and later-generation models reached 86–94% on anesthesiology exams. Task type also matters — accuracy varied by ophthalmology subspecialty, and chart-based questions were excluded because they were unsupported. Reproducibility on clinical pharmacy questions was high (>92%), suggesting stable answers to repeated factual questions, while hallucination remained measurable across benchmarks.21645
How much do you trust ChatGPT's answers for factual or professional information?
Your individual answer is private. Only totals are shown.
Your turn
Have your say
Quick votes, open to everyone. See where you stand the moment you vote. Only totals are ever shown.
How do you feel about this?
No votes yetQuick questions from connected pages
Before you go
What to remember
Try to recall each hidden figure before you reveal it. Remembering, not rereading, is what makes it stick.
Accuracy is task- and version-dependent: reported figures span on an ophthalmology exam with GPT-3.5 to 81.5% with GPT-4 on the Japanese licensing exam and 86–94% for current models on anesthesiology exams.
Hallucination is measurable and persistent: about of ChatGPT responses in HaluEval contained fabricated or unverifiable content, and 11–20.1% of anesthesiology exam responses were hallucinations.
On some narrow tasks ChatGPT performed well: of total knee replacement FAQ responses were accurate, and 79% of clinical pharmacy factual questions were answered correctly versus 66% for pharmacists.
Your reading
0 of 4 chaptersThis answer keeps changing
When new evidence or a better source comes in, this page is updated (it's on version 2, last changed 54 minutes ago). Follow it to be told when that happens.
Ask this Sylo
Still wondering about something?
Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.
More on AI
Everything on AI ›How much energy does ChatGPT use?
How much energy does ChatGPT use per query, and how does that compare with other everyday activities?
Does AI help students learn or make them lazier?
Does using AI tools help students learn more effectively, or does it make them lazier and undermine their learning?
Is using ChatGPT for homework cheating?
Why do AI chatbots make things up?
Is AI dangerous?
Is artificial intelligence dangerous, and what does the evidence say about its risks?
When will we have AGI?
When will we have artificial general intelligence (AGI)?
Behind this page
Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.
Discussion
Sources
Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.
- 1Evaluating the Performance of ChatGPT in OphthalmologyOphthalmology Science (Antaki et al.)Published May 4, 2023Checked Oct 10, 2026
“We tested the accuracy of ChatGPT, a large language model (LLM), in the ophthalmology question-answering space using two popular multiple choice question banks used for the high-stakes Ophthalmic Knowledge Assessment Program (OKAP) exam. The testing sets were of easy-to-moderate difficulty and were diversified, including recall, interpretation, practical and clinical decision-making problems. ChatGPT achieved 55.8% and 42.7% accuracy in the two 260-question simulated exams. Its performance varied across subspecialties, with the best results in general medicine and the worst in neuro-ophthalmology and ophthalmic pathology and intraocular tumors. These results are encouraging but suggest that specialising LLMs through domain-specific pre-training may be necessary to improve their performance in ophthalmic subspecialties.”
- 2Accuracy of ChatGPT on Medical Questions in the National Medical Licensing Examination in Japan: Evaluation StudyJMIR Formative Research (Yanagita et al.)Published Oct 3, 2023Checked Oct 10, 2026
“Of the 400 questions, 292 were analyzed. Questions containing charts, which are not supported by ChatGPT, were excluded. The correct response rate for GPT-4 was 81.5% (237/292), which was significantly higher than the rate for GPT-3.5, 42.8% (125/292). Moreover, GPT-4 surpassed the passing standard (>72%) for the NMLE, indicating its potential as a diagnostic and therapeutic decision aid for physicians. GPT-4 reached the passing standard for the NMLE in Japan, entered in Japanese, although it is limited to written questions. As the accelerated progress in the past few months has shown, the performance of the AI will improve as the large language model continues to learn more, and it may well become a decision support system for medical professionals by providing more accurate information.”
- 3Evaluating the accuracy and relevance of ChatGPT responses to frequently asked questions regarding total knee replacementKnee Surgery and Related Research (Zhang et al.)Published Apr 2, 2024Checked Oct 10, 2026
“Most responses were accurate, while all responses were relevant. Of the 50 FAQs, 44/50 (88%) of ChatGPT responses were classified as accurate, achieving a mean Likert grade of 4.6/5 for factual accuracy. On the other hand, 50/50 (100%) of responses were classified as relevant, achieving a mean Likert grade of 4.9/5 for relevance. ChatGPT performed well in providing accurate and relevant responses to FAQs regarding TKR, demonstrating great potential as a tool for patient education. However, it is not infallible and can occasionally provide inaccurate medical information. Patients and clinicians intending to utilize this technology should be mindful of its limitations and ensure adequate supervision and verification of information provided.”
- 4Performance of ChatGPT on Factual Knowledge Questions Regarding Clinical PharmacyThe Journal of Clinical Pharmacology (Nuland et al.)Published Apr 16, 2024Checked Oct 10, 2026
“Accuracy was defined as the correctness of the answer, and results were compared to the overall score by pharmacists over 2022. Responses were marked concordant if no contradictions were present. The quality of the substantiation was graded by two independent pharmacists using a 4-point scale. Reproducibility was established by presenting questions multiple times and on various days. ChatGPT yielded accurate responses for 79% of the questions, surpassing pharmacists' accuracy of 66%. Concordance was 95%, and the quality of the substantiation was deemed good or excellent for 73% of the questions. Reproducibility was consistently high, both within day and between days (>92%), as well as across different users. ChatGPT demonstrated a higher accuracy and reproducibility to factual knowledge questions related to clinical pharmacy practice than pharmacists. Consequently, we posit that ChatGPT could serve as a valuable resource to pharmacists. We hope the technology will further improve, which may lead to enhanced future performance.”
- 5HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsConference on Empirical Methods in Natural Language Processing (EMNLP) (Li et al.)Published Jan 1, 2023Checked Oct 10, 2026
“Large language models (LLMs), such as Chat-GPT, are prone to generate hallucinations, i.e., content that conflicts with the source or cannot be verified by the factual knowledge.To understand what types of content and to which extent LLMs are apt to hallucinate, we introduce the Hallucination Evaluation benchmark for Large Language Models (HaluEval), a large collection of generated and human-annotated hallucinated samples for evaluating the performance of LLMs in recognizing hallucination.To generate these samples automatically, we propose a two-stage framework, i.e., samplingthen-filtering.Besides, we hire some human labelers to annotate the hallucinations in Chat-GPT responses.The empirical results suggest that ChatGPT is likely to generate hallucinated content related to specific topics by fabricating unverifiable information (i.e., about 19.5% responses).Moreover, existing LLMs face great challenges in recognizing the hallucinations in texts.However, our experiments also prove that providing external knowledge or adding reasoning steps can help LLMs recognize hallucinations.Our benchmark can be accessed at https://github.com/RUCAIBox/HaluEval.”
- 6Performance and Hallucination Analysis of Large Language Models on European Anesthesiology Examinations: Cross-Sectional Comparative Study.JMIR formative research (Andrei et al.)Published Sep 1, 2026Checked Oct 10, 2026
“Statistical analysis included Friedman and Wilcoxon signed-rank tests with Holm-Bonferroni correction, the Cochran Q test, and generalized estimating equations.ResultsAverage success rates ranged from 86% (SD 18%) to 94% (SD 10%) across LLMs and examination types, exceeding the EDAIC part I passing threshold, representing substantial improvement over previously reported GPT-3.5 performance. For the EDAIC, overall intermodel differences were significant (Friedman χ23=13.9; P=.003; W=0.02), with Gemini outperforming GPT-5 as the only pairwise difference. Hallucination rates ranged from 11% (11/100) to 20.1% (44/219) without significant intermodel differences. All models exceeded the EDAIC passing threshold.ConclusionsCurrent-generation LLMs demonstrated consistently high performance across multiple European anesthesiology examinations but continue to produce clinically relevant hallucinations, supporting their role as supervised educational tools rather than autonomous learning resources. These findings underscore the need for structured integration frameworks and systematic verification when deploying LLMs as learning tools in medical education.”
- 7Large language models in healthcare: applications, evaluation frameworks, and governance pathways - a scoping review and multidimensional framework.Frontiers in digital health (Ferreira & Rosa)Published Sep 18, 2026Checked Oct 10, 2026
“The evidence base was dominated by benchmark and simulated-workflow studies (89%), with limited prospective workflow-embedded evaluations (11%). Reported benefits concentrated on documentation efficiency, text quality, and knowledge synthesis; safety-relevant risks included hallucinated content, omission of clinically critical information, demographic bias, privacy vulnerabilities, limited explainability, and automation bias. Studies were geographically concentrated in North America and East Asia, with limited representation from Sub-Saharan Africa, South Asia, and Latin America.ConclusionsCurrent evidence supports cautious deployment of LLMs in selected healthcare tasks under structured oversight. Translational progress depends on prospective evaluation, standardised reporting, equity-focused audits, and lifecycle governance with continuous monitoring. The proposed five-dimensional framework (technical performance, clinical validity, equity, workflow integration, governance) coupled with a three-tier risk model is intended to support researchers and healthcare organisations in assessing readiness and implementing LLM-enabled tools responsibly.”
How it changed
Published 1 time since Oct 10, 2026.
- Version 2Oct 10, 2026Live now
AI-prepared Starting Map from live research.
- First published version.
Help improve it
The brief is open about what's uncertain. These are the specific gaps that new material would fill.
“What affects reliability” has no evidence or firsthand experience yet
It's a synthesis for now. Evidence or experience would show whether it holds.
Open questions
How accurate is ChatGPT in prospective, workflow-embedded clinical use, given that only 11% of the reviewed evidence base consists of such studies?
No answers yet
How do hallucination rates compare across benchmarks and model versions, given figures of about 19.5% (HaluEval) and 11–20.1% (anesthesiology exams) come from different settings?
No answers yet
What are the error rates on non-medical tasks, since the available evaluations concentrate on medical exams, patient FAQs and pharmacy questions?
No answers yet
Do accuracy and reliability hold in settings outside North America and East Asia, which dominate the reviewed evidence?
No answers yet
Around this topic
Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.