Can you trust AI chatbots for medical advice?
Two 2023 studies disagree: chatbots were preferred over physicians on forum questions, yet a medical model still scored below clinicians.
Covers: Large language model chatbots answering medical questions from the public or on medical exams. Clinician-facing decision tools are only mentioned.
3 free full reads left this month. Join or upgrade
The short answer
Evidence-backed AI-prepared starting mapTwo 2023 studies point in different directions. On patient-style questions from a public social media forum, evaluators preferred chatbot responses to physician responses in 78.6% of 585 evaluations, and rated 78.5% of chatbot responses as good or very good quality versus 22.1% of physician responses — about 3.6 times higher prevalence. On medical exams, an instruction-tuned model (Flan-PaLM) reached 67.6% accuracy on MedQA (US Medical Licensing Exam-style questions), surpassing the prior state of the art by more than 17%, but human evaluation found key gaps and the resulting Med-PaLM model remained inferior to clinicians.12
- Evidence 8
- Interpretation 6
In brief
Evaluators preferred chatbot responses to physician responses in 78.6% of 585 evaluations of patient-style forum questions, and rated 78.5% of chatbot answers good or very good versus 22.1% of physician answers.1
Evidence-backedOn medical exam-style questions, Flan-PaLM reached 67.6% accuracy on MedQA, beating the prior state of the art by more than 17%, but the aligned Med-PaLM model remained inferior to clinicians in human evaluation.2
Evidence-backedPreference and quality ratings are not the same as medical correctness or safety; the forum study did not measure health outcomes.1
Interpretation
At a glance
The picture in numbers
Live · updated just now
78.6%
79 in every 100
- Chatbot78.5%
- Physician22.1%
67.6%
68 in every 100
The evidence behind it
6 sources- Other studies and data6
Published in 2023 and 2026
| Source | Kind | Year |
|---|---|---|
| Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum | Other studies and data | 2023 |
| Large language models encode clinical knowledge | Other studies and data | 2023 |
| Cross-sectional comparative evaluation of five large language model-driven chatbots on an expert-curated 44-question set of parent-facing questions about pediatric vitamin D deficiency. | Other studies and data | 2026 |
| Calibration of Self-Reported Confidence and Accuracy of Large Language Models in Medical Question Answering | Other studies and data | 2026 |
| A comparative cross-sectional evaluation of generative AI chatbots for patient-oriented bacterial vaginosis health advice: safety, accuracy, guideline concordance, empathy, and readability. | Other studies and data | 2026 |
| A comparative evaluation of generative AI chatbots for patient-oriented advice on painful diabetic peripheral neuropathy: safety, accuracy, guideline concordance, actionability, and readability. | Other studies and data | 2026 |
The community around it
- Contributions
- 0
- People
- 0
- Following
- 0
Nobody has added anything yet. Experience, evidence or a different view would show up here.
What it means for you
Which fits you?
Pick the situation closest to yours. Each answer says what it rests on.
If you are judging a chatbot's answer by how fluent or empathetic it feels
treat that as a rating of style and perceived quality, not evidence that the advice is medically correct — the forum study measured preference and quality, not outcomes.1
InterpretationIf you are impressed by a chatbot's score on medical exam questions
note that high multiple-choice accuracy coexisted with human evaluation finding key gaps and clinician-superior performance in the same study.2
Evidence-backedIf you are considering using a chatbot to draft replies to patient questions
the forum study's authors suggest this as a use case worth exploring, with physicians editing the drafts, and call for randomized trials before conclusions about outcomes.1
Evidence-backedIf you want to know whether chatbots improve patient outcomes or reduce clinician burnout
that question is not yet answered by these studies; the forum study explicitly frames it as something randomized trials should assess.1
InterpretationThe full story · 2 chapters
01
What the evidence shows
AI summary:One study found chatbot answers preferred and rated higher quality than physician answers; another found a medical model beat exam records but stayed inferior to clinicians.
Evidence-backed: In a cross-sectional study of 195 questions and responses drawn from a public social media forum, evaluators preferred chatbot responses to physician responses in 78.6% of 585 evaluations (95% CI, 75.0%–81.8%). Chatbot responses were longer (mean 211 words, IQR 168–245) than physician responses (mean 52 words, IQR 17–62), and were rated significantly higher on quality. The share rated good or very good (≥4) was 78.5% for the chatbot versus 22.1% for physicians — 3.6 times higher prevalence. The authors note this was a cross-sectional study and call for randomized trials to test whether AI assistants improve responses, reduce clinician burnout, or improve patient outcomes.1
Evidence-backed: On medical knowledge tests, Flan-PaLM achieved state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA, MedMCQA, PubMedQA, and MMLU clinical topics), including 67.6% on MedQA, surpassing the prior state of the art by more than 17%. However, human evaluation revealed key gaps, and the instruction-tuned Med-PaLM model, while encouraging, remained inferior to clinicians. The authors report that comprehension, knowledge recall, and reasoning improve with model scale and instruction prompt tuning, and stress the need for evaluation frameworks and method development to create safe, helpful clinical LLMs.2
Interpretation: Read together, the two studies measure different things: one measures how people rate the style and perceived quality of answers to patient-style questions, the other measures exam accuracy and clinician-judged quality. High preference ratings on a forum do not establish that the advice was medically correct, and high exam scores do not establish that advice is safe in a real conversation.12
Have you asked an AI chatbot a health question?
Your individual response is private. Only totals are shown.
02
Where the debate stands
AI summary:Optimism rests on the forum preference study; caution rests on the exam study where human evaluation found gaps and clinicians still did better.
Interpretation: The case for optimism rests on the forum study: chatbot answers were preferred in most evaluations and rated good or very good far more often than physician answers, and the authors suggest chatbots could draft responses for physicians to edit. The case for caution rests on the exam study: despite record multiple-choice accuracy, human evaluation found key gaps and the aligned model remained inferior to clinicians. A reasonable reading is that current models are strong at producing fluent, well-liked answers and at recalling tested medical knowledge, while clinician judgment still outperforms them where human evaluators look closely.12
Ask this Sylo
Still wondering about something?
Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.
Behind this page
Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.
Discussion
Sources
Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.
- 1Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media ForumJAMA Internal Medicine (Ayers et al.)Published Apr 28, 2023Checked Sep 30, 2026
“Of the 195 questions and responses, evaluators preferred chatbot responses to physician responses in 78.6% (95% CI, 75.0%-81.8%) of the 585 evaluations. Mean (IQR) physician responses were significantly shorter than chatbot responses (52 [17-62] words vs 211 [168-245] words; t = 25.4; P < .001). Chatbot responses were rated of significantly higher quality than physician responses (t = 13.3; P < .001). The proportion of responses rated as good or very good quality (≥ 4), for instance, was higher for chatbot than physicians (chatbot: 78.5%, 95% CI, 72.3%-84.1%; physicians: 22.1%, 95% CI, 16.4%-28.2%;). This amounted to 3.6 times higher prevalence of good or very good quality responses for the chatbot. In this cross-sectional study, a chatbot generated quality and empathetic responses to patient questions posed in an online forum. Further exploration of this technology is warranted in clinical settings, such as using chatbot to draft responses that physicians could then edit. Randomized trials could assess further if using AI assistants might improve responses, lower clinician burnout, and improve patient outcomes.”
- 2Large language models encode clinical knowledgeNature (Singhal et al.)Published Jul 12, 2023Checked Sep 30, 2026
“In addition, we evaluate Pathways Language Model 1 (PaLM, a 540-billion parameter LLM) and its instruction-tuned variant, Flan-PaLM 2 on MultiMedQA. Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA 3 , MedMCQA 4 , PubMedQA 5 and Measuring Massive Multitask Language Understanding (MMLU) clinical topics 6 ), including 67.6% accuracy on MedQA (US Medical Licensing Exam-style questions), surpassing the prior state of the art by more than 17%. However, human evaluation reveals key gaps. To resolve this, we introduce instruction prompt tuning, a parameter-efficient approach for aligning LLMs to new domains using a few exemplars. The resulting model, Med-PaLM, performs encouragingly, but remains inferior to clinicians. We show that comprehension, knowledge recall and reasoning improve with model scale and instruction prompt tuning, suggesting the potential utility of LLMs in medicine. Our human evaluations reveal limitations of today’s models, reinforcing the importance of both evaluation frameworks and method development in creating safe, helpful LLMs for clinical applications.”
- 3A comparative evaluation of generative AI chatbots for patient-oriented advice on painful diabetic peripheral neuropathy: safety, accuracy, guideline concordance, actionability, and readability.Frontiers in public health (Gu et al.)Published Sep 3, 2026Checked Oct 4, 2026
“Inter-rater agreement was high for safety (Fleiss' κ = 0.874), accuracy [ICC (2,1) = 0.881], guideline concordance [ICC (2,1) = 0.874], and actionability [ICC (2,1) = 0.877]. Twenty-eight responses (9.3%) were classified as unsafe or potentially unsafe. Unsafe-response rates ranged from 5.0 to 15.0%, with no detected overall between-model difference (Cochran's Q = 4.462, p = 0.347). Accuracy, guideline concordance, and actionability differed across models (all p W = 0.683, 0.730, and 0.556, respectively). ChatGPT generally achieved higher content-related scores, whereas Doubao scored lower. All readability indices also differed across models (all p ConclusionThe five chatbots showed distinct performance patterns across content quality and readability. Although no overall safety difference was detected, every system generated at least one response with a plausible pathway to inappropriate self-management, delayed assessment, medication or product misuse, or preventable injury. Chatbots may support general patient education, but medication decisions, foot-risk assessment, and urgent-care triage require professional verification.”
- 4A comparative cross-sectional evaluation of generative AI chatbots for patient-oriented bacterial vaginosis health advice: safety, accuracy, guideline concordance, empathy, and readability.Frontiers in reproductive health (Miao et al.)Published Sep 17, 2026Checked Oct 4, 2026
“Observed interface-specific rates ranged from 6.6% (95% CI, 2.6%-15.7%) to 14.8% (95% CI, 8.0%-25.7%), while the matched binary comparison did not detect an overall difference (Cochran's Q = 2.596, df = 4, P = 0.627). Under the documented query conditions, differences were detected in accuracy (Kendall's W = 0.686), study-specific guideline-anchored concordance (W = 0.243), empathy (W = 0.322), and all six readability indices (W range, 0.482-0.679; all P ConclusionsThis exploratory study characterizes a recorded sample of specific chatbot responses, not stable or repeatable performance characteristics. Clinically relevant risks were observed in every interface, and the non-significant safety comparison should not be interpreted as equivalence. Because each researcher-developed prompt was submitted once and the interfaces were queried on different dates in a fixed order, the findings do not establish a stable ranking of underlying model capability, clinical effectiveness, or suitability for unsupervised care.”
- 5Calibration of Self-Reported Confidence and Accuracy of Large Language Models in Medical Question AnsweringJournal of Medical Systems (Boie et al.)Published Jun 26, 2026Checked Oct 4, 2026
“Mean ECE differed by model (Claude Sonnet 4.5 best: 0.06; Gpt-4o worst: 0.127) and varied across specialties ("Skin" best: 0.041; "Social & Preventive Medicine" worst: 0.141). Accuracy of examined LLMs showed analogous variation between specialties. DISCUSSION: Our results demonstrate that high accuracy does not guarantee reliable uncertainty estimation. We identified substantial heterogeneity across medical specialties, where pooled metrics masked a threefold ECE increase between best- and worst-performing domains. We recommend incorporating calibration reporting into LLM evaluations, as larger models exhibit improved "self-knowledge", but uneven overconfidence persists.”
- 6Cross-sectional comparative evaluation of five large language model-driven chatbots on an expert-curated 44-question set of parent-facing questions about pediatric vitamin D deficiency.Frontiers in pediatrics (Shi et al.)Published Sep 9, 2026Checked Oct 4, 2026
“Twenty of 220 responses (9.1%) were unsafe; safety did not differ significantly across models (Cochran's Q = 3.704, df = 4, P = 0.448). Significant inter-model differences were observed in accuracy, empathy, DISCERN, EQIP, JAMA, GQS, and all readability indices. ChatGPT had the highest median accuracy [5.00 (4.00, 5.00)] and DISCERN score [69.60 (65.25, 72.40)]; DeepSeek and Doubao had the highest empathy scores [5.00 (4.80, 5.00)]. ChatGPT and Doubao shared the highest EQIP median (86.00), Doubao had the highest GQS [5.00 (4.00, 5.00)], and Gemini had the highest FRES. JAMA scores were low across models.ConclusionsIn this single-run evaluation, the five chatbot services showed domain-specific performance differences without a significant safety difference. Findings are exploratory rather than evidence of stable model superiority and support multidimensional evaluation with clinician involvement for high-risk or individualized pediatric advice.”
How it changed
Published 1 time since Sep 30, 2026.
- Version 3Sep 30, 2026Live now
Initial Starting Map on the accuracy and safety of LLM chatbots for health questions, built from two 2023 studies: one showing evaluators preferred chatbot answers to physician answers on a public forum, and one showing strong exam accuracy but clinician-superior human evaluation.
- First published version.
Help improve it
The brief is open about what's uncertain. These are the specific gaps that new material would fill.
Open questions
Do chatbot answers to health questions lead to correct understanding and safe decisions in real conversations, rather than only high ratings or exam scores?
No answers yet
What kinds of health questions produce confident but wrong answers, and how often?
No answers yet
How do chatbot answers compare with clinicians on accuracy and safety, not just on perceived quality and empathy?
No answers yet
Would using chatbots to draft responses improve patient outcomes or reduce clinician burnout, as the forum study suggests testing?
No answers yet
Around this topic
Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.