← 返回研究新闻
A teacher and diverse university learners discuss a cyan three-branch overlay linking mobile, conversation and writing activities
综述证据综述20262026年7月26日· 14 min

Impact of artificial intelligence tools on learning motivation in English instruction: A network meta-analysis

Liwei Hsu, Yu-Chun Wang

Asian-Pacific Journal of Second and Foreign Language Education

500 字摘要

A teacher and diverse university learners discuss a cyan three-branch overlay linking mobile, conversation and writing activities

Hsu and Wang synthesize a fast-growing but fragmented literature on whether artificial-intelligence tools strengthen motivation in English instruction. Their 2026 open-access network meta-analysis compares generative-AI chatbots, AI writing assistants, and AI language-learning applications with traditional instruction. The paper is timely for AIEDHK because it reports encouraging effects while also showing why product rankings must be interpreted cautiously: every AI category was compared directly with conventional teaching, but none of the included studies directly compared one AI category with another.

The authors searched Web of Science Core Collection, Scopus, ERIC, PsycINFO, and Google Scholar for peer-reviewed English-language studies published from January 2015 through December 2025. The initial search returned 2,156 records. After duplicate removal, title and abstract screening, and full-text eligibility checks, 16 studies met the criteria. Together they included 1,923 K-12 and university learners, with individual samples ranging from 50 to 412 and a median of 85. Fourteen studies were conducted in university settings and two in K-12 education.

The evidence base was geographically concentrated. Nine studies came from China, three from Iran, and one each from the United Arab Emirates, Algeria, Nigeria, and the United States. Eleven studies examined chatbot or conversational systems, two examined AI language-learning applications, and three examined AI writing assistants. Interventions lasted from six weeks to one semester, with a median duration of eight weeks. Fourteen studies used traditional instruction as the comparison condition.

The review process used independent screening by two reviewers, with Cohen's kappa values of 0.87 for titles and abstracts and 0.92 for full-text eligibility. A quarter of extracted data was double-coded, producing an intraclass correlation of 0.94. The authors assessed randomized trials with the Cochrane RoB 2 tool and quasi-experiments with ROBINS-I. Because blinding is difficult in educational technology studies, many studies had moderate risk in performance-related domains.

Using a frequentist random-effects network model, the authors calculated standardized mean differences as Hedges' g. All three AI categories showed statistically significant positive effects on learning motivation compared with traditional instruction. AI language-learning applications produced the largest pooled estimate, g equals 0.907 with a 95 percent confidence interval from 0.752 to 1.063. Generative-AI chatbots followed at g equals 0.824, with a confidence interval from 0.690 to 0.959. AI writing assistants produced g equals 0.692, with a wider interval from 0.417 to 0.967.

The ranking analysis placed language-learning applications first, chatbots second, and writing assistants third. Yet that order is preliminary. The evidence network was star-shaped: all direct comparisons connected an AI intervention to traditional instruction, so every AI-to-AI comparison was inferred through the common control. Confidence intervals for the pairwise comparisons among AI categories overlapped, and none of those differences was statistically significant. The two language-app studies and three writing-assistant studies also provide much thinner evidence than the eleven chatbot studies.

Several robustness checks were reassuring within those boundaries. Overall heterogeneity was moderate, with I-squared of 42.3 percent. Node-splitting tests did not identify significant inconsistency, and Egger's regression did not indicate significant funnel-plot asymmetry. Removing three studies with elevated risk of bias changed each category's effect by less than 0.06. These checks support the overall finding that AI-supported approaches can improve motivation relative to the included comparison conditions, but they do not turn indirect category rankings into head-to-head evidence.

The study also identifies mechanisms worth testing rather than assuming. Language-learning applications may support competence and autonomy through adaptive difficulty, progress markers, and self-paced practice. Chatbots may reduce anxiety by offering a low-stakes conversational partner. Writing assistants can provide task-specific feedback but may create a more transactional experience. The authors organize these possibilities into an AI-Scaffolded English-Medium Instruction framework: foundational language practice, task-specific writing support, and interactive conversational engagement, with teacher involvement across the levels.

Important limits remain. Measures of motivation varied across studies. The median intervention lasted only eight weeks, so the influence of novelty and the possibility of longer-term motivational decline remain uncertain. Most studies were recent, and initial enthusiasm could inflate effects. The geographic concentration, particularly the nine Chinese studies, limits generalization to other educational systems. Motivation is also not the same as language proficiency, knowledge retention, or independent performance after support is removed.

For Hong Kong schools and universities, the useful conclusion is not to select a product category from the ranking table. A stronger pilot would begin with a specific motivational barrier, choose a tool whose interaction design addresses that barrier, compare it with a credible existing practice, and measure both motivation and independent learning over time. Teacher scaffolding, proficiency differences, equitable access, and intentional fading of assistance should be part of the intervention. The paper supports thoughtful AI-assisted English learning, but it also makes the next research need clear: direct, multi-arm comparisons that test which tools help which learners, under which teaching conditions, and whether the gains persist.

相关论文

大学生向教师解释几何作图,同学在旁思考,桌上的电脑展示相关数字图解
政策 / 伦理2026年9月7日
政策 / 伦理 112

评论:Astra 的 AGI 主张,让教育更需要看见人的真实学习

AIED.HK Editorial

AI Product News Commentary

OpenAI 于 2026 年 9 月 3 日发布 GPT-6 Astra,并引发关于 AGI 是否已经到来的讨论。本文将这一说法视为需要归属的主张,而非已确立的共识。教育眼前的挑战,是分清 AI 能产出什么,以及学习者能独立解释、质疑和迁移什么,再以更强的代理能力支持真正的学习。

产品新闻评论GPT-6 Astra
阅读 500 字摘要 →
三位教育与软件同事在明亮的大学设计工作室审查图解教材卡、注释图表和数字原型
政策 / 伦理2026年9月7日
政策 / 伦理 113

评论:Fable 5.1 把更长程的 AI 工作带进 AIED,教育验证更显重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 于 2026 年 9 月 1 日发布 Claude Fable 5.1,强化长程编程与知识工作能力,并降低缓存读取价格。AIED 的机会,是加快从教学构想到可审查原型与研究分析的循环;真正的考验,是能否把速度转化为更好的教学与可信证据,同时计入总成本、数据条件和人工审核。

产品新闻评论Claude Fable 5.1
阅读 500 字摘要 →
A lecturer and two university students inspect ranked learning tools, separate cloud and local plugin cards, and a review ledger in a bright computing studio
政策 / 伦理2026年8月23日
政策 / 伦理 111

Product news: ChatGPT plugin ranking and Claude Code 2.1.239 make tool selection and workspace boundaries inspectable

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: ChatGPT now ranks plugin recommendations partly by continued use after installation and adds more time-aware answers, while Claude Code 2.1.239 distinguishes cloud-synced plugins from local installations and makes a data-residency cost premium visible. Gemini for Education supplies the institutional purpose boundary across teaching, learning and work. Together, the updates make tool selection, context, cost and human review part of AI workflow literacy.

product newsChatGPT pluginsClaude Code 2.1.239
阅读 500 字摘要 →