
Experimental evidence on the learning impact of generative AI: gains persisted when students used it for explanation rather than automation
Zara Contractor, Germán Reyes
arXiv working paper
Резюме на 500 слов

Contractor and Reyes examine a question that is easy to obscure in product demonstrations: when students have access to a general-purpose generative-AI system while learning a new topic, do they learn more once the system is removed? Their July 2026 working paper reports a randomized experiment conducted in proctored, in-person undergraduate sessions. Participants studied an unfamiliar subject and wrote an analytical essay either with or without access to an off-the-shelf generative-AI tool. They then completed unaided assessments immediately and one week later. The study measures both knowledge tests and open-ended writing, so it separates short-term task performance from later independent learning more clearly than a satisfaction survey can.
The authors report that AI access increased immediate factual and conceptual test scores by 0.27 standard deviations. They also report that the advantage persisted at the one-week assessment. That result matters because a common concern is that AI may improve the visible product while shifting effort away from understanding. In this setting, the reported knowledge gains did not disappear when students worked without the tool. The paper also reports little change in essay quality while AI was available, but better style and relevance in unaided writing one week later.
The most useful finding is not a general claim that AI access is beneficial. The researchers distinguish augmentation-oriented use from automation-oriented use. Students who used the system to obtain explanations of concepts had stronger delayed gains than students who used it primarily to generate text. The paper links the result to reported changes in effort: AI users shifted time away from drafting and toward reading and searching for information, while also reporting greater learning enjoyment. These measures identify plausible mechanisms, but they do not prove every learner followed the same path or that every tool configuration will produce the same effect.
The evidence needs careful interpretation. This is an arXiv working paper rather than a peer-reviewed journal article, and the results should not be generalized without replication. The task involved an unfamiliar topic, proctored in-person sessions and a specific assessment schedule; different courses, age groups, prompting supports, incentives or unrestricted home use may lead to different behavior. The paper's own results suggest that usage quality is central. Giving students a tool without a learning design can encourage either explanation, inquiry and revision or fast drafting with little durable understanding.
For higher education and Hong Kong classrooms, a defensible pilot would make the augmentation route explicit. Teachers can ask students to request explanations, compare them with course sources, annotate what changed in their understanding and complete an independent follow-up task. Rubrics can reward source evaluation, reasoning and revision rather than polished first drafts. Process logs should support reflection rather than surveillance, and assessments should include moments when learners demonstrate what they can do unaided. The study offers a promising but provisional message: AI may support learning when it redirects effort toward sense-making, not when it quietly replaces it.


