
AI lesson agents adapted detectably to learner personas, but delivery lagged content and an LLM judge misranked the leaders
Yi-Cheng Lin, Yu-Kai Guo, Szu-Chi Chen, Bo-Han Feng, Yun-Man Hsu, Hsiang Hsieh, Yu-Jung Lin, Yue-Ling Wu, Jia-Kai Dong, An-Yu Cheng, Yu-Han Huang, Lok-Lam Ieong, Kuan-Yu Chen, Ming-Douo Tchouang, Shao-Hua Sun, Che Lin, Jian-Jiun Ding, Hung-yi Lee
Teaching Monster Challenge 2026 / arXiv preprint
Lin and colleagues benchmarked end-to-end AI instructional-video agents with learner personas across secondary STEM. Seventy-seven teams produced 1,612 preliminary videos: content accuracy and lesson logic were stronger than visual delivery and learner adaptation, 17% of videos received a critical-fact-error flag, and the automated judge's ranking of the ten strongest systems correlated poorly with crowd judgment. Human raters could detect persona adaptation, but the study evaluated teaching-quality proxies rather than learning gains.