
The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems
Kurt VanLehn
Educational Psychologist
৫০০-শব্দের সারাংশ

VanLehn's review is widely cited because it challenges a common assumption in the tutoring literature: that adult human tutoring is dramatically more effective than computer tutoring, and that finer-grained interaction always produces much larger learning gains. The paper reviews experiments comparing human tutoring, computer tutoring, and no tutoring. No tutoring refers to instruction covering the same content without an interactive tutor. Computer tutoring systems are grouped by the granularity of interaction: answer-based systems, step-based systems, and substep-based systems. Most intelligent tutoring systems fall into the step-based or substep-based categories because they can respond during problem solving rather than only after a final answer.
The review found that common beliefs about effect sizes were not well supported by the available evidence. Earlier narratives often suggested that human tutoring had a very large effect, intelligent tutoring systems had a smaller but still large effect, and simpler computer tutoring had a modest effect. VanLehn's synthesis reported lower human tutoring effects than expected and found that intelligent tutoring systems were nearly as effective as human tutoring in the reviewed comparisons. This result made the paper central to debates about the practical value of ITS.
For reviewers, the paper's strength is that it separates evidence from slogans. Human tutoring, intelligent tutoring, and computer-based support are not single interventions; they vary in interaction quality, curriculum fit, feedback timing, domain structure, and implementation. VanLehn's granularity framework gives readers a way to compare systems more carefully. Answer-based systems may only react after a final response, while step-based and substep-based systems can intervene during the reasoning process. That distinction still matters for AI product design. A chatbot that answers questions, a tutor that monitors problem-solving steps, and a dashboard that recommends practice may all be called AI tutoring, but they offer different forms of learning support and should be evaluated differently.
The key implication is not that machines should replace tutors. Rather, the paper reframes what makes tutoring effective. It suggests that well-designed interaction around problem-solving steps can produce substantial learning gains. It also complicates simplistic claims about more granular interaction always being better. The design of feedback, tasks, domain modelling, and student engagement matters. A tutoring system's effectiveness cannot be inferred from whether it uses AI alone.
For AIEDHK, this paper is valuable because it links AI tutoring to evidence of learning impact. It can help the Research News page avoid product-style claims such as AI tutor equals human tutor while still acknowledging that intelligent tutoring systems have a serious empirical basis. The practical message is that AI tutoring should be evaluated through controlled comparisons, effect sizes, and learning outcomes, not demo impressions. In classroom deployment, the best question is where AI can provide scalable practice and feedback, and where human teachers should focus attention. That makes the paper a useful bridge between academic evaluation and procurement-level judgment.


