← 返回研究新聞
Editorial cover for a randomized comparison of structured AI tutoring and active-learning physics classes
期刊論文同行評審研究20252026年7月20日· 9 min

AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting

Greg Kestin, Kelly Miller, Anna Klales, Timothy Milbourne, Gregorio Ponti

Scientific Reports

500 字摘要

Editorial cover for a randomized comparison of structured AI tutoring and active-learning physics classes

Kestin, Miller, Klales, Milbourne, and Ponti test whether a deliberately designed generative-AI tutor can match or exceed a strong classroom comparison rather than a passive lecture. Their open-access Scientific Reports article describes a randomized crossover experiment in a large introductory physics course at Harvard University. The central comparison is between an AI tutor built around established learning principles and an in-class active-learning lesson covering the same content.

The study involved 194 undergraduates and two consecutive lessons on surface tension and fluid flow. Students were divided into two groups. In the first week, one group completed an AI-supported lesson at home while the other attended the active-learning class; the conditions were reversed in the second week. Both groups completed a pre-test and post-test for each topic. The researchers also measured time on task and asked students about engagement, enjoyment, motivation, and growth mindset.

The tutor was not a generic chatbot. Physics instructors created question-specific prompts, structured the interaction to manage cognitive load, embedded content-rich explanations and videos, and required the model to scaffold the learner rather than simply provide answers. This design choice is crucial: the experiment evaluates a pedagogically engineered system operating inside a defined lesson, not unrestricted use of a general-purpose AI assistant.

Students in the AI-tutored condition achieved higher post-test scores. The median post-test score was 4.5 for the AI group and 3.5 for the in-class group, against a combined pre-test median of 2.75. The authors report that median learning gains were more than twice as large with the AI tutor and that the difference was statistically significant. Regression analyses controlling for prior physics knowledge, course performance, ChatGPT experience, topic, test version, and time on task produced a large estimated effect. The authors place the effect between 0.73 and 1.3 standard deviations after addressing a ceiling effect.

The AI condition was also faster for many learners. The classroom lesson provided about 60 minutes of learning time after tests, while the median AI-tutor time on task was 49 minutes; 70 percent of AI users spent less than 60 minutes. Students reported higher engagement in the AI condition and also rated enjoyment, motivation, and growth-mindset-related experience positively. The individualized pace and immediate feedback may help explain both the cognitive and affective results.

The findings need careful boundaries. This was a short intervention covering two physics topics at one selective university. The comparison was not a full course, and the post-tests measured immediate mastery rather than long-term retention, transfer, collaboration, or independent problem solving. The result may depend on GPT-4, expert-written prompts, high-quality instructional videos, a tightly structured framework, and content that fits stepwise tutoring. The authors explicitly caution that the tutor may not outperform classroom active learning for complex synthesis or higher-order critical thinking.

For AIEDHK, the paper offers a productive contrast to studies of unrestricted AI. The result does not show that any chatbot improves learning. It shows that pedagogical architecture can matter as much as model capability. Schools and universities considering AI tutors should document the tutor's instructional rules, compare it with a credible teaching practice, and measure retention and transfer after access is removed. A strong pilot would also examine who benefits, which learners disengage, how errors are handled, and how the tutor complements human discussion. The practical lesson is to evaluate a designed learning system—not merely access to a model.

相關論文

大學生向教師解釋幾何作圖,同學在旁思考,桌上的電腦展示相關數碼圖解
政策 / 倫理2026年9月7日
政策 / 倫理 112

評論:Astra 的 AGI 主張,讓教育更需要看見人的真實學習

AIED.HK Editorial

AI Product News Commentary

OpenAI 於 2026 年 9 月 3 日發布 GPT-6 Astra,並引發關於 AGI 是否已經到來的討論。本文將這一說法視為需要歸屬的主張,而非已確立的共識。教育眼前的挑戰,是分清 AI 能產出甚麼,以及學習者能獨立解釋、質疑和遷移甚麼,再以更強的代理能力支援真正的學習。

產品新聞評論GPT-6 Astra
閱讀 500 字摘要 →
三位教育與軟件同事在明亮的大學設計工作室審查圖解教材卡、註釋圖表和數碼原型
政策 / 倫理2026年9月7日
政策 / 倫理 113

評論:Fable 5.1 把更長程的 AI 工作帶進 AIED,教育驗證更顯重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 於 2026 年 9 月 1 日發布 Claude Fable 5.1,強化長程編程與知識工作能力,並降低快取讀取價格。AIED 的機會,是加快從教學構想到可審查原型與研究分析的循環;真正的考驗,是能否把速度轉化為更好的教學與可信證據,同時計入總成本、資料條件和人工審核。

產品新聞評論Claude Fable 5.1
閱讀 500 字摘要 →
A lecturer and two university students inspect ranked learning tools, separate cloud and local plugin cards, and a review ledger in a bright computing studio
政策 / 倫理2026年8月23日
政策 / 倫理 111

Product news: ChatGPT plugin ranking and Claude Code 2.1.239 make tool selection and workspace boundaries inspectable

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: ChatGPT now ranks plugin recommendations partly by continued use after installation and adds more time-aware answers, while Claude Code 2.1.239 distinguishes cloud-synced plugins from local installations and makes a data-residency cost premium visible. Gemini for Education supplies the institutional purpose boundary across teaching, learning and work. Together, the updates make tool selection, context, cost and human review part of AI workflow literacy.

product newsChatGPT pluginsClaude Code 2.1.239
閱讀 500 字摘要 →