العودة إلى أخبار البحث
A diverse mathematics class works with an AI tutoring system while a teacher compares student reasoning and assessment evidence
مراجعةEvidence synthesis20269 يوليو 2026· 3 min

Review of 12 studies finds promising mathematics-tutoring gains while generative-AI achievement evidence remains thin

Neo Molemane, Moeketsi Mosia, Felix O. Egara

Discover Education

ملخص 500 كلمة

A diverse mathematics class works with an AI tutoring system while a teacher compares student reasoning and assessment evidence

Molemane, Mosia and Egara review empirical evidence on AI-based tutoring and assessment systems in mathematics education. Following PRISMA procedures, they searched Scopus and Web of Science for studies published from 2015 through 2025. The search produced 1,749 records, 76 full texts were assessed, and 12 studies met the final inclusion criteria. Two reviewers independently screened, extracted and appraised the evidence, resolving disagreements through discussion. Methodological quality was examined with the Mixed Methods Appraisal Tool, while findings were synthesized narratively because interventions, learners and outcomes were too heterogeneous for a defensible pooled effect estimate.

Across the included studies, AI tutoring and assessment systems generally showed promising associations with mathematics achievement, engagement or personalized support. Several reports suggested that lower-performing learners could benefit particularly from adaptive guidance, immediate feedback and practice matched to current understanding. The stronger claims came from experimental or quasi-experimental work rather than perception surveys. Effects were not uniform: they varied with learner characteristics, instructional context and implementation fidelity. A system's label as artificial intelligence therefore did not determine its educational value; the design of feedback, teacher integration and opportunities for students to reason remained important.

The review draws a useful distinction between established tutoring systems and recent generative AI. Studies involving ChatGPT or comparable tools mainly reported positive perceptions, usability or engagement. Direct evidence that these tools improved mathematics achievement was still limited. The review consequently calls for larger, more rigorous trials and clearer accounts of intervention mechanisms. This boundary matters because a fluent explanation, a learner's positive rating and a correctly completed exercise are different outcomes. None alone demonstrates durable conceptual understanding, independent problem solving or transfer to an unfamiliar mathematical task.

The evidence base is small and diverse. Twelve studies cannot support a universal estimate across ages, curricula, countries and AI architectures, and the narrative synthesis provides no overall percentage gain. Restricting the search to two databases may have missed eligible work, while publication and language choices can shape what was found. Some included studies were authored by members of the review team, a relationship the paper acknowledges. The journal also identifies the available article as an early-access manuscript that had not completed final editorial processing, so minor presentation or metadata changes could follow.

For Hong Kong, the review supports carefully designed trials rather than immediate system-wide procurement. A school could compare an AI tutor with existing teacher-led practice, map every task to the local curriculum, and examine outcomes separately for learners beginning at different achievement levels. Assessment should include an unaided near-transfer problem, delayed retention and an explanation scored without knowing the student's condition. Chinese and English mathematical language, accessibility, device access, teacher workload and the consistency of feedback across common misconceptions should be recorded alongside scores.

The paper's defensible conclusion is that AI tutoring and assessment in mathematics is promising but conditional, while the achievement evidence for generative AI remains notably thin. Decision-makers should ask which mechanism produced a result, for whom, under what instructional conditions and for how long. Publishing implementation fidelity and null findings would make local evidence more useful than a feature demonstration. Until larger comparative studies accumulate, the review is best used to design stronger evaluations, not to certify every AI mathematics product as effective.

أوراق ذات صلة

Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory
ورقة مجلة2026
ورقة مجلة 54

Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward

Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi

arXiv working paper

In a preregistered randomized online experiment with 1,174 Argentine adults, GPT-4.1 assistance raised workplace-style problem-solving performance for both education groups and reduced the baseline gap from 0.548 to 0.139 standard deviations. Lower-education participants retained a modest gain after AI was removed, but stronger follow-up performance appeared when intensive assistance was paired with sustained human effort.

generative AIrandomized experimenteducation inequality
اقرأ ملخص 500 كلمة
A university student compares an AI explanation with handwritten concept notes while an instructor and peers work in a seminar room
ورقة مجلة2026
ورقة مجلة 50

Experimental evidence on the learning impact of generative AI: gains persisted when students used it for explanation rather than automation

Zara Contractor, Germán Reyes

arXiv working paper

A randomized, proctored experiment reported that undergraduate access to off-the-shelf generative AI raised immediate factual and conceptual test performance by 0.27 standard deviations and that the gains persisted one week later. The working paper also finds a consequential usage pattern: students who used AI to explain concepts showed stronger delayed gains than students who used it to automate drafting.

generative AIrandomized experimenthigher education
اقرأ ملخص 500 كلمة
Chinese secondary students complete homework with digital assistance before taking a separate closed-book assessment observed by a teacher
ورقة مجلة2026
ورقة مجلة 68

Generative AI adoption was linked to higher homework scores but lower unaided exams in a 26,811-student panel

David Strömberg, Victor Lei, Yanhui Wu

CEPR Discussion Paper No. 21577

A CEPR discussion paper analyzes 30 months of records from 26,811 Chinese students in Grades 7–12. Its difference-in-differences estimates associate generative-AI adoption with homework scores 18% higher and completion time 30% lower, but with substantial declines on closed-book and entrance examinations.

generative AIsecondary educationhomework outsourcing
اقرأ ملخص 500 كلمة