← 返回研究新闻
Turkish high-school students solve mathematics problems in three classroom conditions while teachers compare assisted practice with a later unaided test
期刊论文同行评审研究20252026年7月31日· 8 min

Unguarded GPT raised assisted practice scores but reduced unaided mathematics performance in a high-school RCT

Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, Rei Mariman

Proceedings of the National Academy of Sciences

500 字摘要

Turkish high-school students solve mathematics problems in three classroom conditions while teachers compare assisted practice with a later unaided test

Bastani and colleagues test whether the design of a generative-AI tutor changes what students learn, not only what they complete while assisted. Their preregistered cluster randomized controlled trial took place in one large Turkish high school during the 2023–24 school year. About 50 Grade 9, 10 and 11 classes, comprising nearly 1,000 students, completed four 90-minute mathematics sessions. Classes were assigned to ordinary textbooks and notes, a ChatGPT-like GPT-4 interface called GPT Base, or a guarded GPT Tutor.

Each session separated assisted practice from independent performance. Students first solved practice problems with their assigned resources, then completed a related exam without any resource. GPT Base used a general tutor prompt and could provide answers. GPT Tutor was instructed to give hints rather than direct solutions and received teacher-authored correct solutions, common errors and feedback guidance for each problem. That problem-specific information made the tutor more reliable but required substantial teacher preparation.

Both AI conditions improved visible practice performance. Relative to control, GPT Base raised practice scores by 48% and GPT Tutor by 127%. The unaided exam reversed the picture for the unguarded system: GPT Base students scored 17% below control. GPT Tutor students were statistically indistinguishable from control, meaning the guardrails largely removed the penalty but did not produce a significant positive exam gain.

System accuracy and student behaviour help explain the difference. When repeatedly asked for answers to the study's 57 practice problems, GPT Base was correct only 51% of the time, with 42% logical errors and 8% arithmetic errors. Yet the researchers find that errors alone do not explain the exam penalty. Chat logs show that GPT Base students often requested and copied solutions, while GPT Tutor students increasingly attempted answers and asked for help. Students also did not appear to recognise how superficial use could affect later performance.

The study should not be generalized into a claim that all generative AI harms learning. It covers one school, one subject, four sessions and a GPT-4-era configuration. The material represented about 15% of the semester curriculum. The guarded tutor was unusually well supplied with teacher-created problem knowledge and remained reactive rather than diagnosing misconceptions like an expert human tutor. Longer-term retention and transfer were not established.

For classroom design, the decisive outcome is what students can do after assistance is removed. A school can restrict direct answers, require a learner attempt before a hint, ground feedback in verified solutions and include a short unaided transfer problem. Logs can be used for aggregate design improvement rather than punitive surveillance.

For Hong Kong mathematics education, replication should test local curricula, bilingual interaction and varied achievement levels. The result gives a precise warning and a design direction: a fluent tool can raise immediate scores while weakening engagement, and pedagogical guardrails can mitigate harm only when teachers define the knowledge, hints and independent evidence that matter.

Future tutors should be compared on verified accuracy, the quality of elicited reasoning and delayed transfer, not only on practice scores. Teacher workload for building and maintaining guardrails must also be counted when assessing whether the design can scale.

相关论文

A programming lecturer and two diverse university students inspect compiled code, an inheritance diagram, and a grading rubric in a computer laboratory
期刊论文2026
期刊论文 102

Five AI systems outscored the average OOP cohort but still failed compilation and advanced concepts

Marina Lepp, Joosep Kaimre

arXiv preprint

Lepp and Kaimre evaluated ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and Microsoft 365 Copilot on authentic introductory OOP tests and examinations using student grading criteria. Systems exceeded the historical average and often solved long tasks, yet some code did not compile and interfaces, abstract classes, inheritance, and image-based questions remained difficult. The results challenge take-home assessment validity without proving student learning.

programming assessmentobject-oriented programminggenerative AI
阅读 500 字摘要 →
A diverse group of university students explores a branching media-technology learning story while an instructor traces where quiz choices connect to the narrative
会议论文2026
会议论文 90

AI-generated learning stories were clear and well paced, but their quizzes did not belong in the plot

Finn Rogosch, Andreas Schrader

EDULEARN26 Proceedings

Rogosch and Schrader tested AI-generated interactive-fiction episodes with 22 STEM higher-education participants. The five-to-ten-minute stories were rated clear and appropriately long, but story-content coherence averaged below the neutral midpoint and engagement sat near it. Participants most often questioned why characters suddenly demanded technical answers, showing that a playable educational story can still fail to integrate its learning task.

interactive fictioneducational gamesgenerative AI
阅读 500 字摘要 →
Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory
期刊论文2026
期刊论文 54

Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward

Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi

arXiv working paper

In a preregistered randomized online experiment with 1,174 Argentine adults, GPT-4.1 assistance raised workplace-style problem-solving performance for both education groups and reduced the baseline gap from 0.548 to 0.139 standard deviations. Lower-education participants retained a modest gain after AI was removed, but stronger follow-up performance appeared when intensive assistance was paired with sustained human effort.

generative AIrandomized experimenteducation inequality
阅读 500 字摘要 →