SyloSpace

Can you trust AI chatbots for medical advice?

Two 2023 studies disagree: chatbots were preferred over physicians on forum questions, yet a medical model still scored below clinicians.

Updated 4 days ago4 min readVersion 3
CommentsFollow

Covers: Large language model chatbots answering medical questions from the public or on medical exams. Clinician-facing decision tools are only mentioned.

3 free full reads left this month. Join or upgrade

The short answer

Evidence-backed AI-prepared starting map

Two 2023 studies point in different directions. On patient-style questions from a public social media forum, evaluators preferred chatbot responses to physician responses in 78.6% of 585 evaluations, and rated 78.5% of chatbot responses as good or very good quality versus 22.1% of physician responses — about 3.6 times higher prevalence. On medical exams, an instruction-tuned model (Flan-PaLM) reached 67.6% accuracy on MedQA (US Medical Licensing Exam-style questions), surpassing the prior state of the art by more than 17%, but human evaluation found key gaps and the resulting Med-PaLM model remained inferior to clinicians.12

What this rests on2 independent sources
  • Evidence 8
  • Interpretation 6

In brief

  1. Evaluators preferred chatbot responses to physician responses in 78.6% of 585 evaluations of patient-style forum questions, and rated 78.5% of chatbot answers good or very good versus 22.1% of physician answers.1

    Evidence-backed
  2. On medical exam-style questions, Flan-PaLM reached 67.6% accuracy on MedQA, beating the prior state of the art by more than 17%, but the aligned Med-PaLM model remained inferior to clinicians in human evaluation.2

    Evidence-backed
  3. Preference and quality ratings are not the same as medical correctness or safety; the forum study did not measure health outcomes.1

    Interpretation
  4. Both studies are from 2023 and describe models of that period, so their findings may not carry over to later systems.12

    Interpretation

At a glance

The picture in numbers

Live · updated just now

585 evaluations of patient-style forum questions, 2023 study

78.6%

79 in every 100

of evaluations preferred chatbot responses over physician responses1
Patient-style forum questions, 2023 study
  • Chatbot78.5%
  • Physician22.1%
rated good or very good: chatbot answers vs physician answers1
Flan-PaLM model, 2023 study

67.6%

68 in every 100

accuracy on MedQA medical licensing exam-style questions2

The evidence behind it

6 sources
  • Other studies and data6

Published in 2023 and 2026

Sources on this page by kind and year
SourceKindYear
Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media ForumOther studies and data2023
Large language models encode clinical knowledgeOther studies and data2023
Cross-sectional comparative evaluation of five large language model-driven chatbots on an expert-curated 44-question set of parent-facing questions about pediatric vitamin D deficiency.Other studies and data2026
Calibration of Self-Reported Confidence and Accuracy of Large Language Models in Medical Question AnsweringOther studies and data2026
A comparative cross-sectional evaluation of generative AI chatbots for patient-oriented bacterial vaginosis health advice: safety, accuracy, guideline concordance, empathy, and readability.Other studies and data2026
A comparative evaluation of generative AI chatbots for patient-oriented advice on painful diabetic peripheral neuropathy: safety, accuracy, guideline concordance, actionability, and readability.Other studies and data2026

The community around it

Contributions
0
People
0
Following
0

Nobody has added anything yet. Experience, evidence or a different view would show up here.

What it means for you

Which fits you?

Pick the situation closest to yours. Each answer says what it rests on.

If you are judging a chatbot's answer by how fluent or empathetic it feels

treat that as a rating of style and perceived quality, not evidence that the advice is medically correct — the forum study measured preference and quality, not outcomes.1

Interpretation

If you are impressed by a chatbot's score on medical exam questions

note that high multiple-choice accuracy coexisted with human evaluation finding key gaps and clinician-superior performance in the same study.2

Evidence-backed

If you are considering using a chatbot to draft replies to patient questions

the forum study's authors suggest this as a use case worth exploring, with physicians editing the drafts, and call for randomized trials before conclusions about outcomes.1

Evidence-backed

If you want to know whether chatbots improve patient outcomes or reduce clinician burnout

that question is not yet answered by these studies; the forum study explicitly frames it as something randomized trials should assess.1

Interpretation

The full story · 2 chapters

01

What the evidence shows

AI summary:One study found chatbot answers preferred and rated higher quality than physician answers; another found a medical model beat exam records but stayed inferior to clinicians.

Evidence-backed

Evidence-backed: In a cross-sectional study of 195 questions and responses drawn from a public social media forum, evaluators preferred chatbot responses to physician responses in 78.6% of 585 evaluations (95% CI, 75.0%–81.8%). Chatbot responses were longer (mean 211 words, IQR 168–245) than physician responses (mean 52 words, IQR 17–62), and were rated significantly higher on quality. The share rated good or very good (≥4) was 78.5% for the chatbot versus 22.1% for physicians — 3.6 times higher prevalence. The authors note this was a cross-sectional study and call for randomized trials to test whether AI assistants improve responses, reduce clinician burnout, or improve patient outcomes.1

Evidence-backed

Evidence-backed: On medical knowledge tests, Flan-PaLM achieved state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA, MedMCQA, PubMedQA, and MMLU clinical topics), including 67.6% on MedQA, surpassing the prior state of the art by more than 17%. However, human evaluation revealed key gaps, and the instruction-tuned Med-PaLM model, while encouraging, remained inferior to clinicians. The authors report that comprehension, knowledge recall, and reasoning improve with model scale and instruction prompt tuning, and stress the need for evaluation frameworks and method development to create safe, helpful clinical LLMs.2

Interpretation

Interpretation: Read together, the two studies measure different things: one measures how people rate the style and perceived quality of answers to patient-style questions, the other measures exam accuracy and clinician-judged quality. High preference ratings on a forum do not establish that the advice was medically correct, and high exam scores do not establish that advice is safe in a real conversation.12

Participant opinion · poll

Have you asked an AI chatbot a health question?

Have you asked an AI chatbot a health question?Yes, and I acted on itYes, then checked with a professionalYes, just out of curiosityNo
Sign in to respond.

Your individual response is private. Only totals are shown.

02

Where the debate stands

AI summary:Optimism rests on the forum preference study; caution rests on the exam study where human evaluation found gaps and clinicians still did better.

Interpretation

Interpretation: The case for optimism rests on the forum study: chatbot answers were preferred in most evaluations and rated good or very good far more often than physician answers, and the authors suggest chatbots could draft responses for physicians to edit. The case for caution rests on the exam study: despite record multiple-choice accuracy, human evaluation found key gaps and the aligned model remained inferior to clinicians. A reasonable reading is that current models are strong at producing fluent, well-liked answers and at recalling tested medical knowledge, while clinician judgment still outperforms them where human evaluators look closely.12

Evidence-backed

Evidence-backed: Both studies were published in 2023 and describe models of that period; the forum study explicitly calls for randomized trials and further exploration in clinical settings before drawing conclusions about patient outcomes.12

Ask this Sylo

Still wondering about something?

Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.

Behind this page

Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.

Discussion

Nobody has added anything yet. If you have experience, evidence or a different view, you could be the first.

Sources

Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.

  1. 1
    Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum
    JAMA Internal Medicine (Ayers et al.)Published Apr 28, 2023Checked Sep 30, 2026
    “Of the 195 questions and responses, evaluators preferred chatbot responses to physician responses in 78.6% (95% CI, 75.0%-81.8%) of the 585 evaluations. Mean (IQR) physician responses were significantly shorter than chatbot responses (52 [17-62] words vs 211 [168-245] words; t = 25.4; P < .001). Chatbot responses were rated of significantly higher quality than physician responses (t = 13.3; P < .001). The proportion of responses rated as good or very good quality (≥ 4), for instance, was higher for chatbot than physicians (chatbot: 78.5%, 95% CI, 72.3%-84.1%; physicians: 22.1%, 95% CI, 16.4%-28.2%;). This amounted to 3.6 times higher prevalence of good or very good quality responses for the chatbot. In this cross-sectional study, a chatbot generated quality and empathetic responses to patient questions posed in an online forum. Further exploration of this technology is warranted in clinical settings, such as using chatbot to draft responses that physicians could then edit. Randomized trials could assess further if using AI assistants might improve responses, lower clinician burnout, and improve patient outcomes.”
  2. 2
    Large language models encode clinical knowledge
    Nature (Singhal et al.)Published Jul 12, 2023Checked Sep 30, 2026
    “In addition, we evaluate Pathways Language Model 1 (PaLM, a 540-billion parameter LLM) and its instruction-tuned variant, Flan-PaLM 2 on MultiMedQA. Using a combination of prompting strategies, Flan-PaLM achieves state-of-the-art accuracy on every MultiMedQA multiple-choice dataset (MedQA 3 , MedMCQA 4 , PubMedQA 5 and Measuring Massive Multitask Language Understanding (MMLU) clinical topics 6 ), including 67.6% accuracy on MedQA (US Medical Licensing Exam-style questions), surpassing the prior state of the art by more than 17%. However, human evaluation reveals key gaps. To resolve this, we introduce instruction prompt tuning, a parameter-efficient approach for aligning LLMs to new domains using a few exemplars. The resulting model, Med-PaLM, performs encouragingly, but remains inferior to clinicians. We show that comprehension, knowledge recall and reasoning improve with model scale and instruction prompt tuning, suggesting the potential utility of LLMs in medicine. Our human evaluations reveal limitations of today’s models, reinforcing the importance of both evaluation frameworks and method development in creating safe, helpful LLMs for clinical applications.”
  3. 3
    A comparative evaluation of generative AI chatbots for patient-oriented advice on painful diabetic peripheral neuropathy: safety, accuracy, guideline concordance, actionability, and readability.
    Frontiers in public health (Gu et al.)Published Sep 3, 2026Checked Oct 4, 2026
    “Inter-rater agreement was high for safety (Fleiss' κ = 0.874), accuracy [ICC (2,1) = 0.881], guideline concordance [ICC (2,1) = 0.874], and actionability [ICC (2,1) = 0.877]. Twenty-eight responses (9.3%) were classified as unsafe or potentially unsafe. Unsafe-response rates ranged from 5.0 to 15.0%, with no detected overall between-model difference (Cochran's Q = 4.462, p = 0.347). Accuracy, guideline concordance, and actionability differed across models (all p W = 0.683, 0.730, and 0.556, respectively). ChatGPT generally achieved higher content-related scores, whereas Doubao scored lower. All readability indices also differed across models (all p ConclusionThe five chatbots showed distinct performance patterns across content quality and readability. Although no overall safety difference was detected, every system generated at least one response with a plausible pathway to inappropriate self-management, delayed assessment, medication or product misuse, or preventable injury. Chatbots may support general patient education, but medication decisions, foot-risk assessment, and urgent-care triage require professional verification.”
  4. 4
    A comparative cross-sectional evaluation of generative AI chatbots for patient-oriented bacterial vaginosis health advice: safety, accuracy, guideline concordance, empathy, and readability.
    Frontiers in reproductive health (Miao et al.)Published Sep 17, 2026Checked Oct 4, 2026
    “Observed interface-specific rates ranged from 6.6% (95% CI, 2.6%-15.7%) to 14.8% (95% CI, 8.0%-25.7%), while the matched binary comparison did not detect an overall difference (Cochran's Q = 2.596, df = 4, P = 0.627). Under the documented query conditions, differences were detected in accuracy (Kendall's W = 0.686), study-specific guideline-anchored concordance (W = 0.243), empathy (W = 0.322), and all six readability indices (W range, 0.482-0.679; all P ConclusionsThis exploratory study characterizes a recorded sample of specific chatbot responses, not stable or repeatable performance characteristics. Clinically relevant risks were observed in every interface, and the non-significant safety comparison should not be interpreted as equivalence. Because each researcher-developed prompt was submitted once and the interfaces were queried on different dates in a fixed order, the findings do not establish a stable ranking of underlying model capability, clinical effectiveness, or suitability for unsupervised care.”
  5. 5
    Calibration of Self-Reported Confidence and Accuracy of Large Language Models in Medical Question Answering
    Journal of Medical Systems (Boie et al.)Published Jun 26, 2026Checked Oct 4, 2026
    “Mean ECE differed by model (Claude Sonnet 4.5 best: 0.06; Gpt-4o worst: 0.127) and varied across specialties ("Skin" best: 0.041; "Social & Preventive Medicine" worst: 0.141). Accuracy of examined LLMs showed analogous variation between specialties. DISCUSSION: Our results demonstrate that high accuracy does not guarantee reliable uncertainty estimation. We identified substantial heterogeneity across medical specialties, where pooled metrics masked a threefold ECE increase between best- and worst-performing domains. We recommend incorporating calibration reporting into LLM evaluations, as larger models exhibit improved "self-knowledge", but uneven overconfidence persists.”
  6. 6
    Cross-sectional comparative evaluation of five large language model-driven chatbots on an expert-curated 44-question set of parent-facing questions about pediatric vitamin D deficiency.
    Frontiers in pediatrics (Shi et al.)Published Sep 9, 2026Checked Oct 4, 2026
    “Twenty of 220 responses (9.1%) were unsafe; safety did not differ significantly across models (Cochran's Q = 3.704, df = 4, P = 0.448). Significant inter-model differences were observed in accuracy, empathy, DISCERN, EQIP, JAMA, GQS, and all readability indices. ChatGPT had the highest median accuracy [5.00 (4.00, 5.00)] and DISCERN score [69.60 (65.25, 72.40)]; DeepSeek and Doubao had the highest empathy scores [5.00 (4.80, 5.00)]. ChatGPT and Doubao shared the highest EQIP median (86.00), Doubao had the highest GQS [5.00 (4.00, 5.00)], and Gemini had the highest FRES. JAMA scores were low across models.ConclusionsIn this single-run evaluation, the five chatbot services showed domain-specific performance differences without a significant safety difference. Findings are exploratory rather than evidence of stable model superiority and support multidimensional evaluation with clinician involvement for high-risk or individualized pediatric advice.”

How it changed

Published 1 time since Sep 30, 2026.

  1. Version 3Sep 30, 2026Live now

    Initial Starting Map on the accuracy and safety of LLM chatbots for health questions, built from two 2023 studies: one showing evaluators preferred chatbot answers to physician answers on a public forum, and one showing strong exam accuracy but clinician-superior human evaluation.

    • First published version.
Every version, side by side

Help improve it

The brief is open about what's uncertain. These are the specific gaps that new material would fill.

Open questions

  • Do chatbot answers to health questions lead to correct understanding and safe decisions in real conversations, rather than only high ratings or exam scores?

    No answers yet

  • What kinds of health questions produce confident but wrong answers, and how often?

    No answers yet

  • How do chatbot answers compare with clinicians on accuracy and safety, not just on perceived quality and empathy?

    No answers yet

  • Would using chatbots to draft responses improve patient outcomes or reduce clinician burnout, as the forum study suggests testing?

    No answers yet

Around this topic

Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.

Add what you know

Sign in to add what you know. Reading stays open to everyone.

Ask this Sylo

Answers only from “Can you trust AI chatbots for medical advice?”

Ask anything about this page. The AI reads only its reviewed brief, sources and contributions, cites what it used, and says when the page doesn't cover something.