← 返回研究新聞
A diverse secondary-school learning team compares two AI-generated STEM lesson videos against learner personas, visual scaffolds and teacher review notes
工具 / 數據集工具/數據集20262026年8月12日· 9 min

AI lesson agents adapted detectably to learner personas, but delivery lagged content and an LLM judge misranked the leaders

Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen, Bo-Han Feng, Yun-Man Hsu, Hsiang Hsieh, Yu-Jung Lin, Yue-Ling Wu, Jia-Kai Dong, An-Yu Cheng, Yu-Han Huang, Lok-Lam Ieong, Kuan-Yu Chen, Ming-Douo Tchouang, Shao-Hua Sun, Che Lin, Jian-Jiun Ding, Hung-yi Lee

Teaching Monster Challenge 2026 / arXiv preprint

500 字摘要

A diverse secondary-school learning team compares two AI-generated STEM lesson videos against learner personas, visual scaffolds and teacher review notes

Lin and colleagues ask whether an AI agent that generates a complete lesson can also transform subject matter for a particular learner. They frame that capacity as pedagogical content knowledge: representing a concept in a form that fits a learner's age, prior knowledge, attention and gaps. The first Teaching Monster Challenge turns this idea into an end-to-end instructional-video benchmark. Each system receives a course requirement and free-text learner persona and must create one complete video without human intervention.

The benchmark covers Advanced Placement-aligned Physics, Biology, Computer Science and Mathematics at secondary level. Some items hold the topic constant while changing the learner persona, making adaptation observable. Systems were evaluated on content accuracy, pedagogical logic, learner adaptability, and engagement and multimodal presentation. An automated multimodal LLM judge screened submissions, crowd raters compared shortlisted systems blindly, and teachers, school leaders and university professors formed the final expert panel.

The challenge operated at substantial scale. Forty-six teams generated 1,696 warm-up videos. In the preliminary phase, 77 teams generated 1,612 videos for 32 released items. A 22-item automated rubric advanced ten teams; 59 Prolific raters then contributed 246 pairwise comparisons. The top three systems generated 48 videos for 16 final items. Ten experts reviewed the finalists, with each subject assigned two subject specialists and one pedagogy specialist. Award candidates also submitted reproducible system images to verify the no-human-in-the-loop pipeline.

Content accuracy and pedagogical logic scored higher than engagement and learner adaptability. The authors recorded 6,699 negative flags: 39 percent visual delivery, 27 percent learner adaptation, 16 percent content, 8 percent narration and 10 percent count-based explanation errors. Ineffective visual representation appeared in one third of videos; missing scaffolding in 29 percent, jargon overload in 23 percent and prerequisite gaps in 22 percent. About 17 percent received a critical-fact-error flag. These automated flags are not learning outcomes, but they show how a polished lesson can remain overloaded or mismatched.

A separate study tested whether adaptation was perceptible. Sixty-nine Prolific raters contributed 214 ratings after watching a video and choosing its intended learner from three candidates. A persona-independent retrieval baseline stayed at chance, while shortlisted AI systems were identified above it. The human-video comparison was identified more often than the AI group, although the difference was not statistically significant. The agents therefore adapted in detectable ways, but their broader scores suggest that adaptation was often incomplete.

The benchmark also exposes an evaluation problem. Among the ten shortlisted systems, the automated judge's ranking agreed poorly with the crowd ranking: Spearman's rho was -0.17. Scores clustered near the top of the five-point scale. Repeated automated ratings were reasonably stable, so the issue was not merely rerun randomness; the rubric separated a weak tail but did not distinguish leaders as people did. The authors therefore treat automated screening and human comparative judgment as complementary stages.

The boundaries are important. The study evaluates English-language, AP-aligned STEM videos rather than interactive tutoring, classroom implementation or measured learning gains. The final panel mainly reflects Taiwan's education system, and fixed topics, time limits and 2026 systems constrain generalization. The automated judge's flags are not ground truth. For Hong Kong education, the benchmark works best as an audit template: hold a topic constant, vary a bilingual learner profile, inspect whether examples, pacing, prerequisites and visuals truly change, then measure unassisted understanding and transfer with real learners.

相關論文

大學生向教師解釋幾何作圖,同學在旁思考,桌上的電腦展示相關數碼圖解
政策 / 倫理2026年9月7日
政策 / 倫理 112

評論:Astra 的 AGI 主張,讓教育更需要看見人的真實學習

AIED.HK Editorial

AI Product News Commentary

OpenAI 於 2026 年 9 月 3 日發布 GPT-6 Astra,並引發關於 AGI 是否已經到來的討論。本文將這一說法視為需要歸屬的主張,而非已確立的共識。教育眼前的挑戰,是分清 AI 能產出甚麼,以及學習者能獨立解釋、質疑和遷移甚麼,再以更強的代理能力支援真正的學習。

產品新聞評論GPT-6 Astra
閱讀 500 字摘要 →
三位教育與軟件同事在明亮的大學設計工作室審查圖解教材卡、註釋圖表和數碼原型
政策 / 倫理2026年9月7日
政策 / 倫理 113

評論:Fable 5.1 把更長程的 AI 工作帶進 AIED,教育驗證更顯重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 於 2026 年 9 月 1 日發布 Claude Fable 5.1,強化長程編程與知識工作能力,並降低快取讀取價格。AIED 的機會,是加快從教學構想到可審查原型與研究分析的循環;真正的考驗,是能否把速度轉化為更好的教學與可信證據,同時計入總成本、資料條件和人工審核。

產品新聞評論Claude Fable 5.1
閱讀 500 字摘要 →
A lecturer and two university students inspect ranked learning tools, separate cloud and local plugin cards, and a review ledger in a bright computing studio
政策 / 倫理2026年8月23日
政策 / 倫理 111

Product news: ChatGPT plugin ranking and Claude Code 2.1.239 make tool selection and workspace boundaries inspectable

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: ChatGPT now ranks plugin recommendations partly by continued use after installation and adds more time-aware answers, while Claude Code 2.1.239 distinguishes cloud-synced plugins from local installations and makes a data-residency cost premium visible. Gemini for Education supplies the institutional purpose boundary across teaching, learning and work. Together, the updates make tool selection, context, cost and human review part of AI workflow literacy.

product newsChatGPT pluginsClaude Code 2.1.239
閱讀 500 字摘要 →