العودة إلى أخبار البحث
Academic cover for a review comparing human tutoring and intelligent tutoring systems
مراجعةEvidence synthesis20119 يونيو 2026· 2 min

The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems

Kurt VanLehn

Educational Psychologist

ملخص 500 كلمة

Premium tabletop illustration comparing human tutoring, intelligent tutoring systems, noninteractive instruction, feedback granularity, and learning outcome evidence.

VanLehn's review is widely cited because it challenges a common assumption in the tutoring literature: that adult human tutoring is dramatically more effective than computer tutoring, and that finer-grained interaction always produces much larger learning gains. The paper reviews experiments comparing human tutoring, computer tutoring, and no tutoring. No tutoring refers to instruction covering the same content without an interactive tutor. Computer tutoring systems are grouped by the granularity of interaction: answer-based systems, step-based systems, and substep-based systems. Most intelligent tutoring systems fall into the step-based or substep-based categories because they can respond during problem solving rather than only after a final answer.

The review found that common beliefs about effect sizes were not well supported by the available evidence. Earlier narratives often suggested that human tutoring had a very large effect, intelligent tutoring systems had a smaller but still large effect, and simpler computer tutoring had a modest effect. VanLehn's synthesis reported lower human tutoring effects than expected and found that intelligent tutoring systems were nearly as effective as human tutoring in the reviewed comparisons. This result made the paper central to debates about the practical value of ITS.

For reviewers, the paper's strength is that it separates evidence from slogans. Human tutoring, intelligent tutoring, and computer-based support are not single interventions; they vary in interaction quality, curriculum fit, feedback timing, domain structure, and implementation. VanLehn's granularity framework gives readers a way to compare systems more carefully. Answer-based systems may only react after a final response, while step-based and substep-based systems can intervene during the reasoning process. That distinction still matters for AI product design. A chatbot that answers questions, a tutor that monitors problem-solving steps, and a dashboard that recommends practice may all be called AI tutoring, but they offer different forms of learning support and should be evaluated differently.

The key implication is not that machines should replace tutors. Rather, the paper reframes what makes tutoring effective. It suggests that well-designed interaction around problem-solving steps can produce substantial learning gains. It also complicates simplistic claims about more granular interaction always being better. The design of feedback, tasks, domain modelling, and student engagement matters. A tutoring system's effectiveness cannot be inferred from whether it uses AI alone.

For AIEDHK, this paper is valuable because it links AI tutoring to evidence of learning impact. It can help the Research News page avoid product-style claims such as AI tutor equals human tutor while still acknowledging that intelligent tutoring systems have a serious empirical basis. The practical message is that AI tutoring should be evaluated through controlled comparisons, effect sizes, and learning outcomes, not demo impressions. In classroom deployment, the best question is where AI can provide scalable practice and feedback, and where human teachers should focus attention. That makes the paper a useful bridge between academic evaluation and procurement-level judgment.

أوراق ذات صلة

Four diverse researchers compare six separate geometric evidence trays beneath six cyan arrows pointing in different directions
مراجعة2026
مراجعة 46

ChatGPT's impact on student learning outcomes: a meta-analysis of 35 experimental studies

Xinning Wu, Pei Zhu, Jinliang Zhang, Mengwei Yin, Yingxi Wang

Humanities and Social Sciences Communications

A 2026 meta-analysis of 35 experimental studies and 4,193 participants reported a moderate average positive effect of ChatGPT on learning outcomes, but very high heterogeneity means the pooled result should guide conditional design questions rather than a universal effectiveness claim.

ChatGPTmeta-analysislearning outcomes
اقرأ ملخص 500 كلمة
Academic cover for a position paper on ChatGPT and large language models in education
مراجعة2023
مراجعة c3c9681c-ca00-4df2-aa39-490f775df4fb

ChatGPT for good? On opportunities and challenges of large language models for education

Enkelejda Kasneci, Kathrin Sessler, Stefan Kuechemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Guennemann, Eyke Huellermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Juergen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuehn, Gjergji Kasneci

Learning and Individual Differences

A widely cited position paper that balances the educational opportunities of large language models with risks around bias, privacy, assessment, and teacher guidance.

large language modelsChatGPTteacher support
اقرأ ملخص 500 كلمة