
A Large Scale Randomized Control Trial Showing LLM Generated Feedback Helps Low-Knowledge Middle School Math Students with Short-Term Learning
Eamon Worden, Luca Dang, Wen-Chiang Ivan Lim, Sarah Miller, Jiayi Zhang, Aaron Haim, Adam Sales, Ashish Gurung, Neil Heffernan
ACM Learning @ Scale
Résumé de 500 mots

Worden and colleagues provide unusually large causal evidence about LLM-generated feedback in authentic middle-school mathematics practice. Their 2026 ACM Learning @ Scale paper reports two randomized controlled trials embedded in ASSISTments. The first compared AI-generated feedback with business as usual, where students saw only whether an answer was correct. The second compared the same AI-generated approach with feedback previously written by experienced teachers. Across the studies, the authors report a combined 21,478 students working on hundreds of problems, making this one of the largest field evaluations of LLM-generated instructional feedback to date.
The intervention targeted common wrong answers in seventh-grade Illustrative Mathematics fill-in problems. The team first identified incorrect responses submitted by at least 100 students over the preceding five years. Problems containing images were excluded after preliminary tests exposed hallucination risks. Qwen-235B then generated feedback for 3,862 common wrong answers across 653 problems. The goal was not to reveal the solution. Each message acknowledged meaningful progress where possible, identified or responded to the likely error, and offered a concise next-step hint in no more than three sentences.
Human expertise shaped the pipeline before deployment. Two experienced educators advised on the desired task- and process-level feedback, reviewed examples, and helped refine the prompt over five rounds. The final 20-shot prompt covered multiple mathematical skills and misconception patterns. GPT-4.1-mini then served as an automated quality filter for mathematical correctness, grade-level fit, and pedagogical quality. A manual check found a roughly one-percent hallucination rate in the generated set; the identified message and other feedback flagged by the judge were removed. Students also saw a warning that guidance might be AI-written and could be wrong, with a route to report mistakes.
Both trials ran from August 2025 through January 31, 2026. Randomization occurred when a student loaded an eligible problem. The researchers measured two outcomes with mixed-effects logistic regression: whether the learner corrected the current problem on the next attempt, and whether the learner solved the following problem correctly on the first attempt without hints or explanations. The second measure was treated as evidence of near-term transfer or improved self-regulation, although it cannot distinguish those mechanisms.
Study 1 included 87,270 observations from 20,706 unique students, 556 problems, and 1,577 classes. Compared with correctness-only feedback, AI feedback increased the odds of a correct next attempt by 16 percent (odds ratio 1.16, 95 percent confidence interval 1.12 to 1.19). It also increased the odds of first-attempt success on the following problem by 7 percent (odds ratio 1.07, 95 percent confidence interval 1.02 to 1.13). A post-hoc knowledge-group analysis located that transfer benefit among students who had answered zero or one of the previous five problems correctly. Their odds increased by 7 percent, while the medium- and high-knowledge groups showed no detectable benefit.
Study 2 contained 12,950 observations from 6,055 unique students, 97 problems, and 594 classes. The researchers found no statistically discernible difference between AI-generated and teacher-written feedback. For next-attempt correction, the estimated odds ratio was 1.10 but the confidence interval crossed one (0.98 to 1.25). For next-problem success, the estimate was 0.98 with a confidence interval of 0.87 to 1.09. These results are consistent with similar performance, but the smaller sample and wider intervals do not prove strict equivalence. They do show that carefully designed AI feedback was not detectably worse on the measured short-term outcomes.
The study's practical contribution is a human-AI production model rather than an argument for removing teachers. Educators defined the pedagogical form, examples anchored the generation, an automated judge filtered scale output, and the feedback was cached for repeated delivery. The authors estimate that current frontier services could generate a cached message for about five to ten US cents, with the cost amortized when common wrong answers recur. The most frequent wrong answer in the study appeared more than 3,400 times.
Important limits narrow the claim. Per-problem randomization meant many students encountered both conditions, preventing a clean test of sustained exposure. Outcomes covered immediate correction and the next problem, not delayed retention. The study involved one platform, one seventh-grade curriculum, text-only fill-in items, and no demographic data for subgroup analysis. It also combined conceptual misconceptions, procedural errors, and slips. For AIEDHK, the result supports targeted, verified feedback for struggling learners, alongside longer-term assessment, error-type diagnosis, demographic fairness checks, multimodal evaluation, and different feedback styles for students who may benefit more from productive struggle or metacognitive prompts.


