
Impact of artificial intelligence tools on learning motivation in English instruction: A network meta-analysis
Liwei Hsu, Yu-Chun Wang
Asian-Pacific Journal of Second and Foreign Language Education
500-Wörter-Zusammenfassung

Hsu and Wang synthesize a fast-growing but fragmented literature on whether artificial-intelligence tools strengthen motivation in English instruction. Their 2026 open-access network meta-analysis compares generative-AI chatbots, AI writing assistants, and AI language-learning applications with traditional instruction. The paper is timely for AIEDHK because it reports encouraging effects while also showing why product rankings must be interpreted cautiously: every AI category was compared directly with conventional teaching, but none of the included studies directly compared one AI category with another.
The authors searched Web of Science Core Collection, Scopus, ERIC, PsycINFO, and Google Scholar for peer-reviewed English-language studies published from January 2015 through December 2025. The initial search returned 2,156 records. After duplicate removal, title and abstract screening, and full-text eligibility checks, 16 studies met the criteria. Together they included 1,923 K-12 and university learners, with individual samples ranging from 50 to 412 and a median of 85. Fourteen studies were conducted in university settings and two in K-12 education.
The evidence base was geographically concentrated. Nine studies came from China, three from Iran, and one each from the United Arab Emirates, Algeria, Nigeria, and the United States. Eleven studies examined chatbot or conversational systems, two examined AI language-learning applications, and three examined AI writing assistants. Interventions lasted from six weeks to one semester, with a median duration of eight weeks. Fourteen studies used traditional instruction as the comparison condition.
The review process used independent screening by two reviewers, with Cohen's kappa values of 0.87 for titles and abstracts and 0.92 for full-text eligibility. A quarter of extracted data was double-coded, producing an intraclass correlation of 0.94. The authors assessed randomized trials with the Cochrane RoB 2 tool and quasi-experiments with ROBINS-I. Because blinding is difficult in educational technology studies, many studies had moderate risk in performance-related domains.
Using a frequentist random-effects network model, the authors calculated standardized mean differences as Hedges' g. All three AI categories showed statistically significant positive effects on learning motivation compared with traditional instruction. AI language-learning applications produced the largest pooled estimate, g equals 0.907 with a 95 percent confidence interval from 0.752 to 1.063. Generative-AI chatbots followed at g equals 0.824, with a confidence interval from 0.690 to 0.959. AI writing assistants produced g equals 0.692, with a wider interval from 0.417 to 0.967.
The ranking analysis placed language-learning applications first, chatbots second, and writing assistants third. Yet that order is preliminary. The evidence network was star-shaped: all direct comparisons connected an AI intervention to traditional instruction, so every AI-to-AI comparison was inferred through the common control. Confidence intervals for the pairwise comparisons among AI categories overlapped, and none of those differences was statistically significant. The two language-app studies and three writing-assistant studies also provide much thinner evidence than the eleven chatbot studies.
Several robustness checks were reassuring within those boundaries. Overall heterogeneity was moderate, with I-squared of 42.3 percent. Node-splitting tests did not identify significant inconsistency, and Egger's regression did not indicate significant funnel-plot asymmetry. Removing three studies with elevated risk of bias changed each category's effect by less than 0.06. These checks support the overall finding that AI-supported approaches can improve motivation relative to the included comparison conditions, but they do not turn indirect category rankings into head-to-head evidence.
The study also identifies mechanisms worth testing rather than assuming. Language-learning applications may support competence and autonomy through adaptive difficulty, progress markers, and self-paced practice. Chatbots may reduce anxiety by offering a low-stakes conversational partner. Writing assistants can provide task-specific feedback but may create a more transactional experience. The authors organize these possibilities into an AI-Scaffolded English-Medium Instruction framework: foundational language practice, task-specific writing support, and interactive conversational engagement, with teacher involvement across the levels.
Important limits remain. Measures of motivation varied across studies. The median intervention lasted only eight weeks, so the influence of novelty and the possibility of longer-term motivational decline remain uncertain. Most studies were recent, and initial enthusiasm could inflate effects. The geographic concentration, particularly the nine Chinese studies, limits generalization to other educational systems. Motivation is also not the same as language proficiency, knowledge retention, or independent performance after support is removed.
For Hong Kong schools and universities, the useful conclusion is not to select a product category from the ranking table. A stronger pilot would begin with a specific motivational barrier, choose a tool whose interaction design addresses that barrier, compare it with a credible existing practice, and measure both motivation and independent learning over time. Teacher scaffolding, proficiency differences, equitable access, and intentional fading of assistance should be part of the intervention. The paper supports thoughtful AI-assisted English learning, but it also makes the next research need clear: direct, multi-arm comparisons that test which tools help which learners, under which teaching conditions, and whether the gains persist.


