
Unguarded GPT raised assisted practice scores but reduced unaided mathematics performance in a high-school RCT
Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, Rei Mariman
Proceedings of the National Academy of Sciences
৫০০-শব্দের সারাংশ

Bastani and colleagues test whether the design of a generative-AI tutor changes what students learn, not only what they complete while assisted. Their preregistered cluster randomized controlled trial took place in one large Turkish high school during the 2023–24 school year. About 50 Grade 9, 10 and 11 classes, comprising nearly 1,000 students, completed four 90-minute mathematics sessions. Classes were assigned to ordinary textbooks and notes, a ChatGPT-like GPT-4 interface called GPT Base, or a guarded GPT Tutor.
Each session separated assisted practice from independent performance. Students first solved practice problems with their assigned resources, then completed a related exam without any resource. GPT Base used a general tutor prompt and could provide answers. GPT Tutor was instructed to give hints rather than direct solutions and received teacher-authored correct solutions, common errors and feedback guidance for each problem. That problem-specific information made the tutor more reliable but required substantial teacher preparation.
Both AI conditions improved visible practice performance. Relative to control, GPT Base raised practice scores by 48% and GPT Tutor by 127%. The unaided exam reversed the picture for the unguarded system: GPT Base students scored 17% below control. GPT Tutor students were statistically indistinguishable from control, meaning the guardrails largely removed the penalty but did not produce a significant positive exam gain.
System accuracy and student behaviour help explain the difference. When repeatedly asked for answers to the study's 57 practice problems, GPT Base was correct only 51% of the time, with 42% logical errors and 8% arithmetic errors. Yet the researchers find that errors alone do not explain the exam penalty. Chat logs show that GPT Base students often requested and copied solutions, while GPT Tutor students increasingly attempted answers and asked for help. Students also did not appear to recognise how superficial use could affect later performance.
The study should not be generalized into a claim that all generative AI harms learning. It covers one school, one subject, four sessions and a GPT-4-era configuration. The material represented about 15% of the semester curriculum. The guarded tutor was unusually well supplied with teacher-created problem knowledge and remained reactive rather than diagnosing misconceptions like an expert human tutor. Longer-term retention and transfer were not established.
For classroom design, the decisive outcome is what students can do after assistance is removed. A school can restrict direct answers, require a learner attempt before a hint, ground feedback in verified solutions and include a short unaided transfer problem. Logs can be used for aggregate design improvement rather than punitive surveillance.
For Hong Kong mathematics education, replication should test local curricula, bilingual interaction and varied achievement levels. The result gives a precise warning and a design direction: a fluent tool can raise immediate scores while weakening engagement, and pedagogical guardrails can mitigate harm only when teachers define the knowledge, hints and independent evidence that matter.
Future tutors should be compared on verified accuracy, the quality of elicited reasoning and delayed transfer, not only on practice scores. Teacher workload for building and maintaining guardrails must also be counted when assessing whether the design can scale.


