← Retour aux actualités recherche
A diverse mathematics class works with an AI tutoring system while a teacher compares student reasoning and assessment evidence
RevueEvidence synthesis20269 juil. 2026· 3 min

Review of 12 studies finds promising mathematics-tutoring gains while generative-AI achievement evidence remains thin

Neo Molemane, Moeketsi Mosia, Felix O. Egara

Discover Education

Résumé de 500 mots

A diverse mathematics class works with an AI tutoring system while a teacher compares student reasoning and assessment evidence

Molemane, Mosia and Egara review empirical evidence on AI-based tutoring and assessment systems in mathematics education. Following PRISMA procedures, they searched Scopus and Web of Science for studies published from 2015 through 2025. The search produced 1,749 records, 76 full texts were assessed, and 12 studies met the final inclusion criteria. Two reviewers independently screened, extracted and appraised the evidence, resolving disagreements through discussion. Methodological quality was examined with the Mixed Methods Appraisal Tool, while findings were synthesized narratively because interventions, learners and outcomes were too heterogeneous for a defensible pooled effect estimate.

Across the included studies, AI tutoring and assessment systems generally showed promising associations with mathematics achievement, engagement or personalized support. Several reports suggested that lower-performing learners could benefit particularly from adaptive guidance, immediate feedback and practice matched to current understanding. The stronger claims came from experimental or quasi-experimental work rather than perception surveys. Effects were not uniform: they varied with learner characteristics, instructional context and implementation fidelity. A system's label as artificial intelligence therefore did not determine its educational value; the design of feedback, teacher integration and opportunities for students to reason remained important.

The review draws a useful distinction between established tutoring systems and recent generative AI. Studies involving ChatGPT or comparable tools mainly reported positive perceptions, usability or engagement. Direct evidence that these tools improved mathematics achievement was still limited. The review consequently calls for larger, more rigorous trials and clearer accounts of intervention mechanisms. This boundary matters because a fluent explanation, a learner's positive rating and a correctly completed exercise are different outcomes. None alone demonstrates durable conceptual understanding, independent problem solving or transfer to an unfamiliar mathematical task.

The evidence base is small and diverse. Twelve studies cannot support a universal estimate across ages, curricula, countries and AI architectures, and the narrative synthesis provides no overall percentage gain. Restricting the search to two databases may have missed eligible work, while publication and language choices can shape what was found. Some included studies were authored by members of the review team, a relationship the paper acknowledges. The journal also identifies the available article as an early-access manuscript that had not completed final editorial processing, so minor presentation or metadata changes could follow.

For Hong Kong, the review supports carefully designed trials rather than immediate system-wide procurement. A school could compare an AI tutor with existing teacher-led practice, map every task to the local curriculum, and examine outcomes separately for learners beginning at different achievement levels. Assessment should include an unaided near-transfer problem, delayed retention and an explanation scored without knowing the student's condition. Chinese and English mathematical language, accessibility, device access, teacher workload and the consistency of feedback across common misconceptions should be recorded alongside scores.

The paper's defensible conclusion is that AI tutoring and assessment in mathematics is promising but conditional, while the achievement evidence for generative AI remains notably thin. Decision-makers should ask which mechanism produced a result, for whom, under what instructional conditions and for how long. Publishing implementation fidelity and null findings would make local evidence more useful than a feature demonstration. Until larger comparative studies accumulate, the review is best used to design stronger evaluations, not to certify every AI mathematics product as effective.

Articles liés

A programming lecturer and two diverse university students inspect compiled code, an inheritance diagram, and a grading rubric in a computer laboratory
Article de revue2026
Article de revue 102

Five AI systems outscored the average OOP cohort but still failed compilation and advanced concepts

Marina Lepp, Joosep Kaimre

arXiv preprint

Lepp and Kaimre evaluated ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and Microsoft 365 Copilot on authentic introductory OOP tests and examinations using student grading criteria. Systems exceeded the historical average and often solved long tasks, yet some code did not compile and interfaces, abstract classes, inheritance, and image-based questions remained difficult. The results challenge take-home assessment validity without proving student learning.

programming assessmentobject-oriented programminggenerative AI
Lire le résumé de 500 mots →
A diverse group of university students explores a branching media-technology learning story while an instructor traces where quiz choices connect to the narrative
Article de conférence2026
Article de conférence 90

AI-generated learning stories were clear and well paced, but their quizzes did not belong in the plot

Finn Rogosch, Andreas Schrader

EDULEARN26 Proceedings

Rogosch and Schrader tested AI-generated interactive-fiction episodes with 22 STEM higher-education participants. The five-to-ten-minute stories were rated clear and appropriately long, but story-content coherence averaged below the neutral midpoint and engagement sat near it. Participants most often questioned why characters suddenly demanded technical answers, showing that a playable educational story can still fail to integrate its learning task.

interactive fictioneducational gamesgenerative AI
Lire le résumé de 500 mots →
Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory
Article de revue2026
Article de revue 54

Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward

Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi

arXiv working paper

In a preregistered randomized online experiment with 1,174 Argentine adults, GPT-4.1 assistance raised workplace-style problem-solving performance for both education groups and reduced the baseline gap from 0.548 to 0.139 standard deviations. Lower-education participants retained a modest gain after AI was removed, but stronger follow-up performance appeared when intensive assistance was paired with sustained human effort.

generative AIrandomized experimenteducation inequality
Lire le résumé de 500 mots →