
Generative AI adoption was linked to higher homework scores but lower unaided exams in a 26,811-student panel
David Strömberg, Victor Lei, Yanhui Wu
CEPR Discussion Paper No. 21577
Résumé de 500 mots

Strömberg, Lei and Wu examine a pattern that homework dashboards alone can conceal: assisted work may improve while independent performance declines. Their CEPR discussion paper combines 30 months of records for 26,811 Chinese students in Grades 7 through 12 across nine subjects. The data include homework scores and completion time, monthly closed-book examinations and two types of high-stakes entrance examination. Because students and schools adopted generative AI at different times, the authors use a staggered-adoption difference-in-differences design to estimate changes after adoption.
The reported contrast is large. After adoption, homework scores rose by an estimated 18% while homework completion time fell by 30%. Within six months, monthly unaided examination scores were estimated to be 20% lower. Scores on the two high-stakes entrance examinations were estimated to fall by 18% and 24%, with the full entrance-exam penalty appearing after about two years. The paper reports the largest losses in social-science subjects, followed by STEM and languages, and larger estimated effects for younger students, high achievers and boys.
The authors investigate whether use resembles learning support or task outsourcing. About 80% of AI users displayed a combination of unusually high homework scores and unusually short completion time, which the paper interprets as outsourcing-like behaviour. The learning losses concentrate in that group. Users who maintained homework time closer to non-users showed much smaller losses. This pattern is consistent with reduced cognitive engagement, but it is not a direct observation of prompts, copied answers or student intention.
The evidence is consequential but not experimental. AI adoption was not randomly assigned, so the causal interpretation depends on parallel-trends assumptions and on whether other time-varying changes are adequately handled. The outsourcing measure is inferred from administrative traces rather than chat logs. The setting combines one national examination culture and particular homework platforms, and the paper is a discussion paper that may be revised. Its estimates should not be generalized to every learner, subject or guarded tutoring system.
For schools, the study argues for separating product completion from learning evidence. An AI-supported homework task can be paired with a brief unaided explanation, oral defence, retrieval quiz or transfer problem. Teachers can inspect both time-on-task patterns and the reasoning shown, without assuming that fast work is automatically misconduct. Assignments can permit AI for hints, examples and feedback while requiring students to identify what they changed and why.
For Hong Kong education, the strongest lesson is methodological: evaluate the outcome after assistance is removed. Higher homework scores and faster completion may be valuable, but they cannot establish durable learning by themselves. Pilots should compare assisted products, independent assessments and process evidence, while treating the paper's striking estimates as provisional findings from an observational design rather than a universal causal law.
A robust replication would preregister adoption measures and independent outcomes, test the parallel-trends assumption openly, and examine students who use AI for explanation rather than completion. That would help separate harmful outsourcing from forms of support that preserve effort.


