
Review of 12 studies finds promising mathematics-tutoring gains while generative-AI achievement evidence remains thin
Neo Molemane, Moeketsi Mosia, Felix O. Egara
Discover Education
500-Wörter-Zusammenfassung

Molemane, Mosia and Egara review empirical evidence on AI-based tutoring and assessment systems in mathematics education. Following PRISMA procedures, they searched Scopus and Web of Science for studies published from 2015 through 2025. The search produced 1,749 records, 76 full texts were assessed, and 12 studies met the final inclusion criteria. Two reviewers independently screened, extracted and appraised the evidence, resolving disagreements through discussion. Methodological quality was examined with the Mixed Methods Appraisal Tool, while findings were synthesized narratively because interventions, learners and outcomes were too heterogeneous for a defensible pooled effect estimate.
Across the included studies, AI tutoring and assessment systems generally showed promising associations with mathematics achievement, engagement or personalized support. Several reports suggested that lower-performing learners could benefit particularly from adaptive guidance, immediate feedback and practice matched to current understanding. The stronger claims came from experimental or quasi-experimental work rather than perception surveys. Effects were not uniform: they varied with learner characteristics, instructional context and implementation fidelity. A system's label as artificial intelligence therefore did not determine its educational value; the design of feedback, teacher integration and opportunities for students to reason remained important.
The review draws a useful distinction between established tutoring systems and recent generative AI. Studies involving ChatGPT or comparable tools mainly reported positive perceptions, usability or engagement. Direct evidence that these tools improved mathematics achievement was still limited. The review consequently calls for larger, more rigorous trials and clearer accounts of intervention mechanisms. This boundary matters because a fluent explanation, a learner's positive rating and a correctly completed exercise are different outcomes. None alone demonstrates durable conceptual understanding, independent problem solving or transfer to an unfamiliar mathematical task.
The evidence base is small and diverse. Twelve studies cannot support a universal estimate across ages, curricula, countries and AI architectures, and the narrative synthesis provides no overall percentage gain. Restricting the search to two databases may have missed eligible work, while publication and language choices can shape what was found. Some included studies were authored by members of the review team, a relationship the paper acknowledges. The journal also identifies the available article as an early-access manuscript that had not completed final editorial processing, so minor presentation or metadata changes could follow.
For Hong Kong, the review supports carefully designed trials rather than immediate system-wide procurement. A school could compare an AI tutor with existing teacher-led practice, map every task to the local curriculum, and examine outcomes separately for learners beginning at different achievement levels. Assessment should include an unaided near-transfer problem, delayed retention and an explanation scored without knowing the student's condition. Chinese and English mathematical language, accessibility, device access, teacher workload and the consistency of feedback across common misconceptions should be recorded alongside scores.
The paper's defensible conclusion is that AI tutoring and assessment in mathematics is promising but conditional, while the achievement evidence for generative AI remains notably thin. Decision-makers should ask which mechanism produced a result, for whom, under what instructional conditions and for how long. Publishing implementation fidelity and null findings would make local evidence more useful than a feature demonstration. Until larger comparative studies accumulate, the review is best used to design stronger evaluations, not to certify every AI mathematics product as effective.


