← 返回研究新聞
A human tutor guides a middle-school learner through mathematics while a discreet AI panel suggests questions rather than answers
期刊論文同行評審研究20252026年8月1日· 8 min

Tutor CoPilot improved short-term topic mastery by supporting human tutors, with the largest gain among lower-rated tutors

Rose E. Wang, Ana T. Ribeiro, Carly D. Robinson, Susanna Loeb, Dora Demszky

arXiv preprint

500 字摘要

A human tutor guides a middle-school learner through mathematics while a discreet AI panel suggests questions rather than answers

Wang and colleagues test a human-AI design that supports tutors during instruction rather than placing a chatbot in front of students. Tutor CoPilot reads the current mathematics problem and tutor-student conversation, then suggests guidance such as a probing question, explanation, hint or similar problem. The human tutor decides whether to use or edit the suggestion. The preregistered tutor-level randomized trial was conducted with a virtual tutoring provider and nine Title I schools in one southern U.S. district.

At launch, 782 tutors were active, with 386 assigned access and 396 in control. They served 1,787 students in Grades 3 through 8 across 4,136 mathematics sessions over two months, producing more than 550,000 messages. The primary causal outcome was whether a student mastered the lesson topic on the provider's exit ticket. Access to Tutor CoPilot increased that mastery rate by four percentage points in the intention-to-treat analysis. For students taught by tutors who had lower prior quality ratings, the gain was nine percentage points.

Conversation analysis suggests a mechanism. Tutors with access used more guiding questions, asked learners to explain reasoning and gave fewer direct answers. The system may therefore distribute elements of expert tutoring practice at the moment they are needed. At the study's usage level, the authors estimate language-model cost at about 20 dollars per tutor per year. That figure excludes product development, training, monitoring, privacy, support and wider deployment costs.

The longer-term result is essential: the study did not find a statistically significant improvement on the end-of-year mathematics test. The exit ticket is a meaningful proximal measure, but it cannot establish durable achievement or transfer. The work also took place with one provider and district over two months, with variable tool use. Some suggestions were not grade-appropriate, and de-identifying names cannot prevent learners from revealing other personal details in free conversation.

For implementation, the human-in-the-loop architecture is promising only if the human remains capable and accountable. Tutors need training to reject a plausible but poor suggestion, adapt language to the learner and protect personal information. Providers can evaluate suggestion quality by topic, grade and tutor experience, sample interactions for human review and show tutors why a recommended move may help. A system that silently optimizes acceptance rates could undermine that judgment.

For Hong Kong tutoring and school support, a pilot should test local curricula, Cantonese and English dialogue and varied learner needs. Outcomes can include immediate mastery, delayed unaided problems, tutor practice and equity across tutor experience. Full cost should include onboarding and review.

The paper's strongest finding is appropriately narrow: real-time AI support improved short-term topic mastery and especially helped learners served by lower-rated tutors. Whether that support produces sustained learning remains unproven, making delayed independent assessment a necessary next step.

The human tutor should also be able to report a poor suggestion without interrupting the lesson, and programme leaders should study rejection as useful evidence rather than failure. Good oversight learns from the moments when professional judgment overrides the model.

相關論文

大學生向教師解釋幾何作圖,同學在旁思考,桌上的電腦展示相關數碼圖解
政策 / 倫理2026年9月7日
政策 / 倫理 112

評論:Astra 的 AGI 主張,讓教育更需要看見人的真實學習

AIED.HK Editorial

AI Product News Commentary

OpenAI 於 2026 年 9 月 3 日發布 GPT-6 Astra,並引發關於 AGI 是否已經到來的討論。本文將這一說法視為需要歸屬的主張,而非已確立的共識。教育眼前的挑戰,是分清 AI 能產出甚麼,以及學習者能獨立解釋、質疑和遷移甚麼,再以更強的代理能力支援真正的學習。

產品新聞評論GPT-6 Astra
閱讀 500 字摘要 →
三位教育與軟件同事在明亮的大學設計工作室審查圖解教材卡、註釋圖表和數碼原型
政策 / 倫理2026年9月7日
政策 / 倫理 113

評論:Fable 5.1 把更長程的 AI 工作帶進 AIED,教育驗證更顯重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 於 2026 年 9 月 1 日發布 Claude Fable 5.1,強化長程編程與知識工作能力,並降低快取讀取價格。AIED 的機會,是加快從教學構想到可審查原型與研究分析的循環;真正的考驗,是能否把速度轉化為更好的教學與可信證據,同時計入總成本、資料條件和人工審核。

產品新聞評論Claude Fable 5.1
閱讀 500 字摘要 →
A lecturer and two university students inspect ranked learning tools, separate cloud and local plugin cards, and a review ledger in a bright computing studio
政策 / 倫理2026年8月23日
政策 / 倫理 111

Product news: ChatGPT plugin ranking and Claude Code 2.1.239 make tool selection and workspace boundaries inspectable

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: ChatGPT now ranks plugin recommendations partly by continued use after installation and adds more time-aware answers, while Claude Code 2.1.239 distinguishes cloud-synced plugins from local installations and makes a data-residency cost premium visible. Gemini for Education supplies the institutional purpose boundary across teaching, learning and work. Together, the updates make tool selection, context, cost and human review part of AI workflow literacy.

product newsChatGPT pluginsClaude Code 2.1.239
閱讀 500 字摘要 →