← 研究ニュースに戻る
Chinese secondary students complete homework with digital assistance before taking a separate closed-book assessment observed by a teacher
ジャーナル論文Peer-reviewed study20262026年8月6日· 8 min

Generative AI adoption was linked to higher homework scores but lower unaided exams in a 26,811-student panel

David Strömberg, Victor Lei, Yanhui Wu

CEPR Discussion Paper No. 21577

500語要約

Chinese secondary students complete homework with digital assistance before taking a separate closed-book assessment observed by a teacher

Strömberg, Lei and Wu examine a pattern that homework dashboards alone can conceal: assisted work may improve while independent performance declines. Their CEPR discussion paper combines 30 months of records for 26,811 Chinese students in Grades 7 through 12 across nine subjects. The data include homework scores and completion time, monthly closed-book examinations and two types of high-stakes entrance examination. Because students and schools adopted generative AI at different times, the authors use a staggered-adoption difference-in-differences design to estimate changes after adoption.

The reported contrast is large. After adoption, homework scores rose by an estimated 18% while homework completion time fell by 30%. Within six months, monthly unaided examination scores were estimated to be 20% lower. Scores on the two high-stakes entrance examinations were estimated to fall by 18% and 24%, with the full entrance-exam penalty appearing after about two years. The paper reports the largest losses in social-science subjects, followed by STEM and languages, and larger estimated effects for younger students, high achievers and boys.

The authors investigate whether use resembles learning support or task outsourcing. About 80% of AI users displayed a combination of unusually high homework scores and unusually short completion time, which the paper interprets as outsourcing-like behaviour. The learning losses concentrate in that group. Users who maintained homework time closer to non-users showed much smaller losses. This pattern is consistent with reduced cognitive engagement, but it is not a direct observation of prompts, copied answers or student intention.

The evidence is consequential but not experimental. AI adoption was not randomly assigned, so the causal interpretation depends on parallel-trends assumptions and on whether other time-varying changes are adequately handled. The outsourcing measure is inferred from administrative traces rather than chat logs. The setting combines one national examination culture and particular homework platforms, and the paper is a discussion paper that may be revised. Its estimates should not be generalized to every learner, subject or guarded tutoring system.

For schools, the study argues for separating product completion from learning evidence. An AI-supported homework task can be paired with a brief unaided explanation, oral defence, retrieval quiz or transfer problem. Teachers can inspect both time-on-task patterns and the reasoning shown, without assuming that fast work is automatically misconduct. Assignments can permit AI for hints, examples and feedback while requiring students to identify what they changed and why.

For Hong Kong education, the strongest lesson is methodological: evaluate the outcome after assistance is removed. Higher homework scores and faster completion may be valuable, but they cannot establish durable learning by themselves. Pilots should compare assisted products, independent assessments and process evidence, while treating the paper's striking estimates as provisional findings from an observational design rather than a universal causal law.

A robust replication would preregister adoption measures and independent outcomes, test the parallel-trends assumption openly, and examine students who use AI for explanation rather than completion. That would help separate harmful outsourcing from forms of support that preserve effort.

関連論文

A programming lecturer and two diverse university students inspect compiled code, an inheritance diagram, and a grading rubric in a computer laboratory
ジャーナル論文2026
ジャーナル論文 102

Five AI systems outscored the average OOP cohort but still failed compilation and advanced concepts

Marina Lepp, Joosep Kaimre

arXiv preprint

Lepp and Kaimre evaluated ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and Microsoft 365 Copilot on authentic introductory OOP tests and examinations using student grading criteria. Systems exceeded the historical average and often solved long tasks, yet some code did not compile and interfaces, abstract classes, inheritance, and image-based questions remained difficult. The results challenge take-home assessment validity without proving student learning.

programming assessmentobject-oriented programminggenerative AI
500語要約を読む →
A diverse group of university students explores a branching media-technology learning story while an instructor traces where quiz choices connect to the narrative
会議論文2026
会議論文 90

AI-generated learning stories were clear and well paced, but their quizzes did not belong in the plot

Finn Rogosch, Andreas Schrader

EDULEARN26 Proceedings

Rogosch and Schrader tested AI-generated interactive-fiction episodes with 22 STEM higher-education participants. The five-to-ten-minute stories were rated clear and appropriately long, but story-content coherence averaged below the neutral midpoint and engagement sat near it. Participants most often questioned why characters suddenly demanded technical answers, showing that a playable educational story can still fail to integrate its learning task.

interactive fictioneducational gamesgenerative AI
500語要約を読む →
Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory
ジャーナル論文2026
ジャーナル論文 54

Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward

Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi

arXiv working paper

In a preregistered randomized online experiment with 1,174 Argentine adults, GPT-4.1 assistance raised workplace-style problem-solving performance for both education groups and reduced the baseline gap from 0.548 to 0.139 standard deviations. Lower-education participants retained a modest gain after AI was removed, but stronger follow-up performance appeared when intensive assistance was paired with sustained human effort.

generative AIrandomized experimenteducation inequality
500語要約を読む →