
ChatGPT's impact on student learning outcomes: a meta-analysis of 35 experimental studies
Xinning Wu, Pei Zhu, Jinliang Zhang, Mengwei Yin, Yingxi Wang
Humanities and Social Sciences Communications
Ringkasan 500 kata

Wu and colleagues synthesize experimental evidence on ChatGPT and student learning outcomes. Their 2026 open-access meta-analysis includes 35 studies published from 2022 through 2024, 4,193 participants, and 134 effect sizes. The paper asks not only whether ChatGPT-supported instruction is associated with better outcomes, but also whether subject, duration, education level, instructional mode, and knowledge type help explain why results differ.
The authors searched Chinese and English databases for experimental or quasi-experimental studies that compared a ChatGPT-supported condition with another instructional condition and reported sufficient quantitative data. Twenty-eight included studies were written in English and seven in Chinese. Two researchers independently coded study features, with a reported kappa of 0.851. A seven-item quality checklist covered design, active controls, pre-post measures, instrument reliability, baseline equivalence, and clarity of methods and interventions; the average quality score was 6.11 out of seven.
Using a random-effects model and Hedges' g, the pooled effect on student learning outcomes was 0.670, with a 95 percent confidence interval from 0.495 to 0.844. The authors interpret this as a moderate positive effect. Only three of the 35 studies had negative effect sizes. When outcomes were divided into categories, the pooled estimate was 0.872 for cognitive outcomes and 0.539 for non-cognitive outcomes.
The headline average conceals substantial variation. The overall Q statistic was 409.067 and I-squared was 91.444 percent, indicating that most observed variation reflected real differences among studies rather than sampling error alone. A pooled number under this level of heterogeneity should not be read as the result any classroom can expect. It is better treated as a prompt to examine instructional conditions, tasks, measures, and implementation.
Moderator analyses found significant differences by subject, experimental duration, and instructional mode. Education level and the distinction between declarative and procedural knowledge were not significant moderators in the authors' coding. These results suggest that how ChatGPT is embedded in teaching may matter more than simply providing access. However, moderator categories combine diverse interventions, and observational comparisons between study subgroups do not have the causal strength of random assignment within one experiment.
The authors used a funnel plot, fail-safe analysis, and Begg's test and reported no significant publication bias. That is reassuring but not definitive, especially in a new field with rapidly changing products, many small studies, and heterogeneous outcomes. The search window ended in 2024, so the synthesis describes earlier versions of ChatGPT and early adoption practices rather than the full 2026 product environment.
Outcome classification also deserves caution. The cognitive category combines achievement, problem solving, creativity, critical thinking, and social skills, while the non-cognitive category combines engagement, interest, motivation, and self-efficacy. These constructs are not interchangeable. A strong average for a broad category does not establish durable knowledge, independent performance, or higher-order reasoning in every subject.
For Hong Kong schools and universities, the review supports moving beyond the question, “Does ChatGPT work?” A stronger evaluation specifies the learning goal, comparison practice, teacher role, duration, allowed assistance, and independent post-test. It should record fidelity, examine subgroups, and test whether gains persist when ChatGPT is removed. The meta-analysis is encouraging evidence that well-designed uses can help, but its very high heterogeneity makes context and pedagogy the central finding.


