العودة إلى أخبار البحث
Turkish high-school students solve mathematics problems in three classroom conditions while teachers compare assisted practice with a later unaided test
ورقة مجلةPeer-reviewed study202531 يوليو 2026· 3 min

Unguarded GPT raised assisted practice scores but reduced unaided mathematics performance in a high-school RCT

Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, Rei Mariman

Proceedings of the National Academy of Sciences

ملخص 500 كلمة

Turkish high-school students solve mathematics problems in three classroom conditions while teachers compare assisted practice with a later unaided test

Bastani and colleagues test whether the design of a generative-AI tutor changes what students learn, not only what they complete while assisted. Their preregistered cluster randomized controlled trial took place in one large Turkish high school during the 2023–24 school year. About 50 Grade 9, 10 and 11 classes, comprising nearly 1,000 students, completed four 90-minute mathematics sessions. Classes were assigned to ordinary textbooks and notes, a ChatGPT-like GPT-4 interface called GPT Base, or a guarded GPT Tutor.

Each session separated assisted practice from independent performance. Students first solved practice problems with their assigned resources, then completed a related exam without any resource. GPT Base used a general tutor prompt and could provide answers. GPT Tutor was instructed to give hints rather than direct solutions and received teacher-authored correct solutions, common errors and feedback guidance for each problem. That problem-specific information made the tutor more reliable but required substantial teacher preparation.

Both AI conditions improved visible practice performance. Relative to control, GPT Base raised practice scores by 48% and GPT Tutor by 127%. The unaided exam reversed the picture for the unguarded system: GPT Base students scored 17% below control. GPT Tutor students were statistically indistinguishable from control, meaning the guardrails largely removed the penalty but did not produce a significant positive exam gain.

System accuracy and student behaviour help explain the difference. When repeatedly asked for answers to the study's 57 practice problems, GPT Base was correct only 51% of the time, with 42% logical errors and 8% arithmetic errors. Yet the researchers find that errors alone do not explain the exam penalty. Chat logs show that GPT Base students often requested and copied solutions, while GPT Tutor students increasingly attempted answers and asked for help. Students also did not appear to recognise how superficial use could affect later performance.

The study should not be generalized into a claim that all generative AI harms learning. It covers one school, one subject, four sessions and a GPT-4-era configuration. The material represented about 15% of the semester curriculum. The guarded tutor was unusually well supplied with teacher-created problem knowledge and remained reactive rather than diagnosing misconceptions like an expert human tutor. Longer-term retention and transfer were not established.

For classroom design, the decisive outcome is what students can do after assistance is removed. A school can restrict direct answers, require a learner attempt before a hint, ground feedback in verified solutions and include a short unaided transfer problem. Logs can be used for aggregate design improvement rather than punitive surveillance.

For Hong Kong mathematics education, replication should test local curricula, bilingual interaction and varied achievement levels. The result gives a precise warning and a design direction: a fluent tool can raise immediate scores while weakening engagement, and pedagogical guardrails can mitigate harm only when teachers define the knowledge, hints and independent evidence that matter.

Future tutors should be compared on verified accuracy, the quality of elicited reasoning and delayed transfer, not only on practice scores. Teacher workload for building and maintaining guardrails must also be counted when assessing whether the design can scale.

أوراق ذات صلة

Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory
ورقة مجلة2026
ورقة مجلة 54

Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward

Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi

arXiv working paper

In a preregistered randomized online experiment with 1,174 Argentine adults, GPT-4.1 assistance raised workplace-style problem-solving performance for both education groups and reduced the baseline gap from 0.548 to 0.139 standard deviations. Lower-education participants retained a modest gain after AI was removed, but stronger follow-up performance appeared when intensive assistance was paired with sustained human effort.

generative AIrandomized experimenteducation inequality
اقرأ ملخص 500 كلمة
A university student compares an AI explanation with handwritten concept notes while an instructor and peers work in a seminar room
ورقة مجلة2026
ورقة مجلة 50

Experimental evidence on the learning impact of generative AI: gains persisted when students used it for explanation rather than automation

Zara Contractor, Germán Reyes

arXiv working paper

A randomized, proctored experiment reported that undergraduate access to off-the-shelf generative AI raised immediate factual and conceptual test performance by 0.27 standard deviations and that the gains persisted one week later. The working paper also finds a consequential usage pattern: students who used AI to explain concepts showed stronger delayed gains than students who used it to automate drafting.

generative AIrandomized experimenthigher education
اقرأ ملخص 500 كلمة
Chinese secondary students complete homework with digital assistance before taking a separate closed-book assessment observed by a teacher
ورقة مجلة2026
ورقة مجلة 68

Generative AI adoption was linked to higher homework scores but lower unaided exams in a 26,811-student panel

David Strömberg, Victor Lei, Yanhui Wu

CEPR Discussion Paper No. 21577

A CEPR discussion paper analyzes 30 months of records from 26,811 Chinese students in Grades 7–12. Its difference-in-differences estimates associate generative-AI adoption with homework scores 18% higher and completion time 30% lower, but with substantial declines on closed-book and entrance examinations.

generative AIsecondary educationhomework outsourcing
اقرأ ملخص 500 كلمة