Does AI tutoring improve learning?
AI-supported instruction usually matches or slightly beats traditional teaching, but the advantage varies by context and is not always positive.
Covers: This page reviews evidence from randomized controlled trials, meta-analyses, and large-scale studies on AI tutoring systems (e.g., intelligent tutoring systems, chatbots) and their effects on learning outcomes in K-12 and higher education. It does not cover AI tools for administrative tasks or non-educational uses.
3 free full reads left this month. Join or upgrade
The short answer
Evidence-backed AI-organised, reviewedAcross the available evidence, AI-supported instruction generally matches or modestly outperforms traditional instruction on measured learning outcomes, but the size and meaning of the advantage vary widely by context, and results are not uniformly positive. A meta-analysis of AI-supported STEM instruction found a statistically significant overall positive effect (g = 0.67, 95% CI [0.49, 0.85]). In simulated surgical skills, AI tutoring was comparable to expert instructors, with only a small, low-certainty advantage on expert-rated OSATS scores (MD 0.20; 95% CI, 0.01 to 0.39). In nursing anatomy, an AI-enhanced blended model produced higher module scores (82.5 ± 7.3) than blended (77.6 ± 8.1) or traditional teaching (74.1 ± 8.5), and in basic medical pathology an AI-plus-knowledge-graph model produced higher final scores (84.2 ± 6.9 vs. 74.1 ± 7.6; adjusted mean difference 8.2, 95% CI 4.7–11.7; d = 1.13). However, a systematic review of health science education found null, mixed and negative findings alongside favorable ones, particularly when AI was compared with established educational resources or used without structured instructional guidance.12345
- Evidence 23
- Interpretation 3
In brief
AI-supported instruction shows a statistically significant positive overall effect on STEM learning (g = 0.67, 95% CI [0.49, 0.85]), with effectiveness varying by level, discipline, and duration.1
Evidence-backedIn simulated surgical skills, AI tutoring was comparable to expert instructors, with only a small, low-certainty OSATS advantage (MD 0.20) below meaningful competency cut-points, and higher extraneous cognitive load.2
Evidence-backedA non-randomized medical student study found higher knowledge retention with AI-assisted learning at one week (75.5% vs. 57.9%) and one month (61.4% vs. 35.9%), but this association does not establish causation.6
Evidence-backedResults are not uniformly positive: a health science review reports null, mixed and negative findings alongside favorable ones, especially when AI is compared with established resources or used without structured guidance.5
Evidence-backed
At a glance
The picture in numbers
Live · updated just now
- AI + knowledge graph84.2 score
- Traditional teaching74.1 score
- AI-enhanced82.5 score
- Blended77.6 score
- Traditional74.1 score
268 participants
The evidence behind it
8 sources- Reviews of many studies4
- Trials1
- Other studies and data2
- Background1
Published in 2026
| Source | Kind | Year |
|---|---|---|
| The Impact of Artificial Intelligence-Supported Instruction on Student Learning in STEM: A Systematic Review and Meta-Analysis. | Reviews of many studies | 2026 |
| Comparison of artificial intelligence assisted training and traditional learning paths in clinical simulation skills training: meta-analysis of randomized controlled trials. | Reviews of many studies | 2026 |
| Artificial intelligence augmented tutoring vs expert instruction on learning simulated general surgical skills: a systematic review and meta-analysis. | Reviews of many studies | 2026 |
| Artificial Intelligence-Assisted Versus Traditional Learning and Long-Term Knowledge Retention Among Undergraduate Medical Students: A Sequential, Explanatory Mixed-Methods Study. | Other studies and data | 2026 |
| Effectiveness of a generative AI-powered digital tutor integrated with a knowledge graph in anatomy education for nursing students: a randomized controlled trial. | Trials | 2026 |
| Intelligent tutoring system (Wikipedia) | Background | Unknown |
| Effects of Artificial Intelligence-Supported Education on Critical Thinking, Problem-Solving, Clinical Reasoning, and Decision-Making in Health Science Education: A Systematic Review of Interventional Studies. | Reviews of many studies | 2026 |
| Application and effectiveness analysis of AI Intelligent tutoring system combined with knowledge graph in basic medical pathology teaching. | Other studies and data | 2026 |
The community around it
- Contributions
- 0
- People
- 0
- Following
- 0
Nobody has added anything yet. Experience, evidence or a different view would show up here.
What it means for you
Which fits you?
Pick the situation closest to yours. Each answer says what it rests on.
If you are choosing between AI tutoring and expert human instruction for procedural skills
the evidence suggests they are broadly comparable, so a hybrid model may be the better choice, though this needs validation.2
Evidence-backedIf you are teaching complex cognitive skills like clinical reasoning
use AI to replace inefficient peer practice with standardized, high-intensity training and let expert teachers focus on advanced judgment and emotional communication.7
Evidence-backedIf you are designing a STEM course with AI support
expect a moderate positive effect on learning outcomes, with results varying by educational level, discipline, and intervention duration.1
Evidence-backedIf you want students to retain knowledge longer
encourage active engagement with personalized, interactive AI explanations rather than passive reliance on AI-generated answers.6
Evidence-backedIf you are considering replacing human instructors with AI
the available evidence does not support replacement and instead points to hybrid integration.27
Evidence-backedIf you are evaluating AI tutoring for K-12 or non-STEM subjects
the current evidence is concentrated in medical and STEM education, so benefits in other settings are not yet established.12
InterpretationIf you are introducing AI into a health science curriculum
pair it with structured instructional guidance and keep learners actively reasoning, because null, mixed and negative results were reported particularly when AI was used without structured guidance or compared with established resources.5
Evidence-backedIf you are judging whether AI tutoring is worth adopting on the strength of score gains alone
note that gains are measured mostly on immediate or course-level assessments, and validated measures of sustained, unaided performance are still called for.5
Evidence-backedThe full story · 2 chapters
01
What the evidence shows
AI summary:Meta-analyses and course studies mostly show small to moderate learning gains with AI support, though some results are uncertain or not uniformly positive.
Evidence-backed: A systematic review and meta-analysis of AI-supported instruction in STEM education found a statistically significant overall positive effect on student learning outcomes (g = 0.67, 95% CI [0.49, 0.85]). The authors report that effectiveness varied by educational level, STEM discipline, and intervention duration, and that publication-bias analyses suggested minimal influence on the overall findings. The largest effects appeared in interventions lasting more than one month and up to two months, though no consistent pattern of increasing effectiveness with longer duration was found.1
Evidence-backed: In simulated general surgical skills, a meta-analysis of four studies (three RCTs and one pilot prospective study, 268 participants total) found AI tutoring produced a small, statistically significant improvement in expert-rated OSATS scores compared with expert instruction (MD 0.20; 95% CI, 0.01 to 0.39; no heterogeneity, I2 = 0%), but with low certainty. The AI group reported significantly higher extraneous cognitive load (MD 0.23; p = 0.01), and there was no significant difference in ICEMS scores. The authors note the 0.20-point OSATS advantage is of uncertain clinical significance, falling below commonly published competency cut-points, and conclude the findings do not support replacing human instructors with AI.2
Evidence-backed: A randomized controlled trial in nursing anatomy compared an AI-enhanced group using a generative AI-powered digital tutor integrated with a knowledge graph against a blended group and a traditional teaching group, all following the same 16-week course structure. The AI-Enhanced Group achieved higher module scores (82.5 ± 7.3) than the Blended Group (77.6 ± 8.1) and the Traditional Teaching Group (74.1 ± 8.5) (F = 28.64), with significant differences across multiple learning outcomes including knowledge retention and clinical reasoning.3
Evidence-backed: In basic medical pathology teaching, a study comparing an "AI + knowledge graph" integrated model with traditional teaching found the experimental group scored higher on final scores (84.2 ± 6.9 vs. 74.1 ± 7.6); after adjusting for class-level clustering in a two-level mixed-effects model, the adjusted mean difference was 8.2 (95% CI 4.7–11.7, P = 0.012), with Cohen's d = 1.13. The experimental group also showed a higher knowledge mastery rate (90.3% vs. 75.8%) and higher clinical case analysis scores (28.5 ± 3.2 vs. 21.3 ± 4.1).4
Evidence-backed: A sequential mixed-methods study among undergraduate medical students found both AI-assisted and traditional learning significantly improved immediate post-test scores, but AI-assisted learning was associated with significantly higher knowledge retention at one week (75.5% vs. 57.9%) and one month (61.4% vs. 35.9%) (p<0.001). A significantly greater proportion of students exceeded a Bloom-inspired heuristic benchmark for high memory retention after AI-assisted learning (43.4% vs. 3.0%). The authors caution that the non-randomized design means this association does not establish that AI-assisted learning caused the improvement, and that passive reliance on AI-generated responses was associated with cognitive offloading.6
Evidence-backed: A systematic review of interventional studies in health science education examined effects on critical thinking, problem-solving, clinical reasoning and decision-making across conventional and generative AI technologies, including diagnostic AI systems, intelligent tutoring systems, AI-enabled chatbots, generative image systems and large language models. Findings varied across the four outcomes: structured AI-supported interventions frequently showed favorable effects on selected clinical reasoning, decision-making, problem-solving and critical-thinking outcomes, but null, mixed and negative findings were also reported, particularly when AI was compared with established educational resources or used without structured instructional guidance. The authors conclude the evidence is inconsistent across technologies, instructional designs, comparators and cognitive constructs, and that structured integration preserving active learner engagement and independent reasoning appears important.5
Evidence-backed: In clinical simulation skills training, a meta-analysis of RCTs found procedure time and operation completion rate did not differ between AI-assisted and traditional groups, while learner satisfaction, willingness to recommend the training, and stress level were higher in the AI-assisted group. The authors conclude AI is best used to replace inefficient peer practice with standardized, high-intensity training, with expert teachers focusing on advanced clinical judgment, emotional communication, and personalized correction, and that the best path is deep integration with traditional teaching rather than replacing teachers.7
Evidence-backed: Intelligent tutoring systems are computer systems that imitate human tutors to provide immediate, customized instruction or feedback, usually without a human teacher. They typically aim to replicate the demonstrated benefits of one-to-one personalized tutoring in contexts where students would otherwise receive one-to-many classroom instruction or no teacher at all, such as online homework.8
In your experience, how does AI tutoring compare with traditional instruction for your learning?
Your individual response is private. Only totals are shown.
02
Where AI tutoring seems to fit best
AI summary:The evidence favors hybrid models over replacement, with AI best for procedural skills and structured use, and weakest against established resources.
Interpretation: The pattern across these studies points toward hybrid models rather than replacement. The surgical-skills meta-analysis explicitly states the evidence supports a hybrid model rather than replacing human instructors, though it notes this itself requires rigorous empirical validation. The clinical simulation review similarly recommends deep integration with traditional teaching, combining their strengths.27
Evidence-backed: The type of skill appears to matter. The clinical simulation review distinguishes procedural skills, where AI (especially vision-based deep learning) can effectively guide training, from complex cognitive skills like clinical reasoning, where AI is best used to replace inefficient peer practice with standardized, high-intensity training while expert teachers handle advanced judgment and emotional communication.7
Evidence-backed: How students use AI may matter as much as whether they use it. The retention study links students' perceptions of personalized, interactive explanations and active engagement to improved retention, while passive reliance on AI-generated responses was associated with cognitive offloading. The health science review likewise reports that null, mixed and negative findings appeared particularly when AI was used without structured instructional guidance, and that structured integration preserving active engagement and independent reasoning appears important.65
Interpretation: Where AI is compared against established educational resources rather than against no support at all, the advantage is least reliable: the health science review reports null, mixed and negative findings in exactly those comparisons. This suggests the comparator matters as much as the technology.5
Ask this Sylo
Still wondering about something?
Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.
Behind this page
Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.
Discussion
Sources
Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.
- 1The Impact of Artificial Intelligence-Supported Instruction on Student Learning in STEM: A Systematic Review and Meta-Analysis.Journal of Intelligence (Doğan et al.)Published Jun 15, 2026Checked Oct 3, 2026
“Moderator analyses were conducted for educational level, STEM discipline, and intervention duration. Publication bias was assessed using multiple diagnostic methods. The meta-analysis revealed a statistically significant overall positive effect of AI-supported instruction on student learning outcomes in STEM education (g = 0.67, 95% CI [0.49, 0.85], p p = 0.088). Regarding intervention duration, the highest effect size was observed in interventions lasting more than one month and up to two months, though no consistent pattern of increasing effectiveness with longer durations was found. Publication bias analyses suggested minimal influence on the overall findings. AI-supported instructional interventions demonstrate a moderately to highly positive impact on student learning outcomes in STEM education. The effectiveness of these interventions varies according to educational level, disciplinary context, and intervention duration. These findings provide robust empirical evidence supporting the pedagogical value of AI in STEM education and offer guidance for educators and policymakers regarding effective implementation.”
- 2Artificial intelligence augmented tutoring vs expert instruction on learning simulated general surgical skills: a systematic review and meta-analysis.BMC medical education (Hanna et al.)Published Jun 10, 2026Checked Oct 3, 2026
“Four studies (3 RCTs and 1 pilot prospective study), encompassing a total of 268 participants, were included in the meta-analysis. AI tutoring showed a small, statistically significant improvement in expert-rated OSATS scores (MD 0.20; 95% CI, 0.01 to 0.39) with no heterogeneity (I2 = 0%) but with low certainty. The AI group reported a significantly higher extraneous cognitive load (MD 0.23; p = 0.01). No significant difference was found in ICEMS scores.ConclusionAI tutoring systems demonstrated comparable effectiveness to expert instructors in simulated surgical skill acquisition. The small OSATS advantage (0.20 points) is of uncertain clinical significance, falling below commonly published competency cut-points for meaningful change on global rating scales, and is based on low-certainty evidence driven by a single high-risk-of-bias study. AI tutoring imposed a higher extraneous cognitive load. These findings do not support replacing human instructors with AI. Instead, the evidence supports a hybrid model, though this itself requires rigorous empirical validation.”
- 3Effectiveness of a generative AI-powered digital tutor integrated with a knowledge graph in anatomy education for nursing students: a randomized controlled trial.BMC medical education (Zhao et al.)Published May 22, 2026Checked Oct 3, 2026
“All groups followed the same 16-week course structure based on Gagné's Nine Events of Instruction, differing only in the teaching tools and tutoring strategies used. Learning outcomes were evaluated using module examination scores, final comprehensive scores, knowledge graph comprehension scores, knowledge retention rates, case-based inference accuracy, and multidimensional questionnaires. Statistical analyses were conducted using one-way ANOVA with post-hoc comparisons.ResultsSignificant differences were observed among the three groups across multiple learning outcomes. The AI-Enhanced Group achieved higher module scores (82.5 ± 7.3) compared with the Blended Group (77.6 ± 8.1) and Traditional Teaching Group (74.1 ± 8.5) (F = 28.64, p ConclusionsThe blended teaching model integrating a generative AI-powered digital tutor with a knowledge graph significantly improved nursing students' academic performance, knowledge retention, and clinical reasoning ability. This AI-enhanced instructional model may provide an effective approach for reforming foundational medical education and promoting structured knowledge learning.”
- 4Application and effectiveness analysis of AI Intelligent tutoring system combined with knowledge graph in basic medical pathology teaching.BMC medical education (Yuanyuan & Ying)Published May 5, 2026Checked Oct 4, 2026
“The control group used traditional teaching, while the experimental group adopted the "AI + knowledge graph" integrated model with AI-assisted learning and knowledge graph reviews. SPSS was used for quantitative and qualitative data analysis.ResultsThe experimental group showed significantly better outcomes: final score (84.2 ± 6.9 vs. 74.1 ± 7.6), after adjusting for class-level clustering using a two-level mixed-effects model, adjusted mean difference = 8.2, 95% CI: 4.7-11.7, P = 0.012, Cohen's d = 1.13 (95% CI not computed), indicating a large effect size. knowledge mastery rate (90.3% vs.75.8%), clinical case analysis score (28.5±3.2 vs. 21.3±4.1, PConclusionThis integrated model effectively promotes structured knowledge construction and learning efficiency, realizing the synergistic effect of AI and knowledge graph. It provides a replicable and practical reference for the digital and intelligent reform of basic medical education courses.”
- 5Effects of Artificial Intelligence-Supported Education on Critical Thinking, Problem-Solving, Clinical Reasoning, and Decision-Making in Health Science Education: A Systematic Review of Interventional Studies.Advances in medical education and practice (Alhur et al.)Published Sep 8, 2026Checked Oct 4, 2026
“The included studies evaluated conventional and generative AI technologies, including diagnostic AI systems, intelligent tutoring systems, AI-enabled chatbots, generative image systems, and large language models. Findings varied across the four cognitive outcomes. Structured AI-supported interventions frequently demonstrated favorable effects on selected clinical reasoning, decision-making, problem-solving, and critical-thinking outcomes; however, null, mixed, and negative findings were also reported, particularly when AI was compared with established educational resources or used without structured instructional guidance.ConclusionAI-supported education may enhance selected higher-order cognitive outcomes among health science learners, but the evidence is inconsistent across technologies, instructional designs, comparators, and cognitive constructs. Structured integration that preserves active learner engagement and independent reasoning appears important. Further rigorous studies using validated outcome measures and assessments of sustained, unaided performance are needed.”
- 6Artificial Intelligence-Assisted Versus Traditional Learning and Long-Term Knowledge Retention Among Undergraduate Medical Students: A Sequential, Explanatory Mixed-Methods Study.Cureus (Padmavathi et al.)Published Aug 11, 2026Checked Oct 3, 2026
“Both learning methods significantly improved immediate post-test scores. However, AI-assisted learning was associated with significantly higher knowledge retention at 1 week (75.5% vs. 57.9%) and 1 month (61.4% vs. 35.9%) compared with traditional learning (p<0.001). Repeated-measures ANOVA showed significant effects of learning method (F=181.66, p<0.001), assessment time (F=625.73, p<0.001), and their interaction (F=62.23, p<0.001). A significantly greater proportion of students exceeded a Bloom-inspired heuristic benchmark for high memory retention following AI-assisted learning (43.4% vs. 3.0%). AI-assisted learning was associated with higher long-term memory retention compared with traditional learning under the specific, non-randomized conditions investigated; this association does not establish that AI-assisted learning caused the improvement. Students' perceptions of personalized, interactive explanations and active engagement appeared linked to improved retention, whereas passive reliance on AI-generated responses was associated with cognitive offloading.”
- 7Comparison of artificial intelligence assisted training and traditional learning paths in clinical simulation skills training: meta-analysis of randomized controlled trials.Frontiers in medicine (Che et al.)Published Aug 21, 2026Checked Oct 3, 2026
“Procedure time and Operation completion rate did not differ between groups. Learner satisfaction and willingness to recommend the training and stress level of AI assisted training were higher than those of the control group.ConclusionThis study shows that AI, especially the vision-based deep-learning, can effectively guide the training of procedural skills. However, for complex cognitive skills like clinical reasoning, AI is best used to replace inefficient peer practice with standardized, high-intensity training. Expert teachers should then focus on advanced clinical judgment, emotional communication, and personalized correction. Overall, AI offers a major opportunity to improve quality and scale up medical simulation education. Its best path is deep integration with traditional teaching-combining their strengths rather than simply replacing teachers.Systematic review registrationhttps://www.crd.york.ac.uk/PROSPERO/view/CRD420261411402. The protocol was registered in advance in PROSPERO online platform as CRD420261411402.”
- 8Intelligent tutoring system (Wikipedia)WikipediaPublished Sep 30, 2026Checked Oct 3, 2026
“An intelligent tutoring system (ITS) is a computer system that imitates human tutors and aims to provide immediate and customized instruction or feedback to learners, usually without requiring intervention from a human teacher. ITSs have the common goal of enabling learning in a meaningful and effective manner by using a variety of computing technologies. There are many examples of ITSs being used in both formal education and professional settings in which they have demonstrated their capabilities and limitations. There is a close relationship between intelligent tutoring, cognitive learning theories and design; and there is ongoing research to improve the effectiveness of ITS. An ITS typically aims to replicate the demonstrated benefits of one-to-one, personalized tutoring, in contexts where students would otherwise have access to one-to-many instruction from a single teacher (e.g., classroom lectures), or no teacher at all (e.g., online homework). ITSs are often designed with the goal of providing access to high quality education to each and every student.”
How it changed
Published 2 times since Oct 3, 2026.
- Version 3Oct 4, 2026Live now
Added two newly available sources: a systematic review of AI-supported education and higher-order cognitive outcomes in health science education, which reports inconsistent results (favorable, null, mixed and negative findings), and a pathology teaching study reporting a large adjusted effect (8.2 points, d = 1.13).
- The main finding was rewritten.
- Updated “What the evidence shows”.
- Updated “Where AI tutoring seems to fit best”.
- Version 2Oct 3, 2026
AI-prepared Starting Map from live research.
- First published version.
Help improve it
The brief is open about what's uncertain. These are the specific gaps that new material would fill.
Open questions
How does AI tutoring compare with no tutoring at all, rather than with traditional instruction?
No answers yet
Do the benefits seen in medical and STEM education hold in K-12 settings and in non-STEM subjects?
No answers yet
Do the retention advantages observed at one week and one month persist over longer periods?
No answers yet
Why did AI tutoring impose higher extraneous cognitive load in surgical simulation, and can that be reduced without losing its benefits?
No answers yet
What rigorous evidence exists that hybrid AI-plus-teacher models outperform either approach alone?
No answers yet
Which technologies, instructional designs and comparators produce the null, mixed and negative results reported in health science education, and why?
No answers yet
Do AI-supported gains hold up on validated measures of sustained, unaided performance rather than immediate post-tests?
No answers yet
Around this topic
Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.