← 返回研究新聞
A teacher and diverse university learners discuss a cyan three-branch overlay linking mobile, conversation and writing activities
綜述證據綜述20262026年7月26日· 14 min

Impact of artificial intelligence tools on learning motivation in English instruction: A network meta-analysis

Liwei Hsu, Yu-Chun Wang

Asian-Pacific Journal of Second and Foreign Language Education

500 字摘要

A teacher and diverse university learners discuss a cyan three-branch overlay linking mobile, conversation and writing activities

Hsu and Wang synthesize a fast-growing but fragmented literature on whether artificial-intelligence tools strengthen motivation in English instruction. Their 2026 open-access network meta-analysis compares generative-AI chatbots, AI writing assistants, and AI language-learning applications with traditional instruction. The paper is timely for AIEDHK because it reports encouraging effects while also showing why product rankings must be interpreted cautiously: every AI category was compared directly with conventional teaching, but none of the included studies directly compared one AI category with another.

The authors searched Web of Science Core Collection, Scopus, ERIC, PsycINFO, and Google Scholar for peer-reviewed English-language studies published from January 2015 through December 2025. The initial search returned 2,156 records. After duplicate removal, title and abstract screening, and full-text eligibility checks, 16 studies met the criteria. Together they included 1,923 K-12 and university learners, with individual samples ranging from 50 to 412 and a median of 85. Fourteen studies were conducted in university settings and two in K-12 education.

The evidence base was geographically concentrated. Nine studies came from China, three from Iran, and one each from the United Arab Emirates, Algeria, Nigeria, and the United States. Eleven studies examined chatbot or conversational systems, two examined AI language-learning applications, and three examined AI writing assistants. Interventions lasted from six weeks to one semester, with a median duration of eight weeks. Fourteen studies used traditional instruction as the comparison condition.

The review process used independent screening by two reviewers, with Cohen's kappa values of 0.87 for titles and abstracts and 0.92 for full-text eligibility. A quarter of extracted data was double-coded, producing an intraclass correlation of 0.94. The authors assessed randomized trials with the Cochrane RoB 2 tool and quasi-experiments with ROBINS-I. Because blinding is difficult in educational technology studies, many studies had moderate risk in performance-related domains.

Using a frequentist random-effects network model, the authors calculated standardized mean differences as Hedges' g. All three AI categories showed statistically significant positive effects on learning motivation compared with traditional instruction. AI language-learning applications produced the largest pooled estimate, g equals 0.907 with a 95 percent confidence interval from 0.752 to 1.063. Generative-AI chatbots followed at g equals 0.824, with a confidence interval from 0.690 to 0.959. AI writing assistants produced g equals 0.692, with a wider interval from 0.417 to 0.967.

The ranking analysis placed language-learning applications first, chatbots second, and writing assistants third. Yet that order is preliminary. The evidence network was star-shaped: all direct comparisons connected an AI intervention to traditional instruction, so every AI-to-AI comparison was inferred through the common control. Confidence intervals for the pairwise comparisons among AI categories overlapped, and none of those differences was statistically significant. The two language-app studies and three writing-assistant studies also provide much thinner evidence than the eleven chatbot studies.

Several robustness checks were reassuring within those boundaries. Overall heterogeneity was moderate, with I-squared of 42.3 percent. Node-splitting tests did not identify significant inconsistency, and Egger's regression did not indicate significant funnel-plot asymmetry. Removing three studies with elevated risk of bias changed each category's effect by less than 0.06. These checks support the overall finding that AI-supported approaches can improve motivation relative to the included comparison conditions, but they do not turn indirect category rankings into head-to-head evidence.

The study also identifies mechanisms worth testing rather than assuming. Language-learning applications may support competence and autonomy through adaptive difficulty, progress markers, and self-paced practice. Chatbots may reduce anxiety by offering a low-stakes conversational partner. Writing assistants can provide task-specific feedback but may create a more transactional experience. The authors organize these possibilities into an AI-Scaffolded English-Medium Instruction framework: foundational language practice, task-specific writing support, and interactive conversational engagement, with teacher involvement across the levels.

Important limits remain. Measures of motivation varied across studies. The median intervention lasted only eight weeks, so the influence of novelty and the possibility of longer-term motivational decline remain uncertain. Most studies were recent, and initial enthusiasm could inflate effects. The geographic concentration, particularly the nine Chinese studies, limits generalization to other educational systems. Motivation is also not the same as language proficiency, knowledge retention, or independent performance after support is removed.

For Hong Kong schools and universities, the useful conclusion is not to select a product category from the ranking table. A stronger pilot would begin with a specific motivational barrier, choose a tool whose interaction design addresses that barrier, compare it with a credible existing practice, and measure both motivation and independent learning over time. Teacher scaffolding, proficiency differences, equitable access, and intentional fading of assistance should be part of the intervention. The paper supports thoughtful AI-assisted English learning, but it also makes the next research need clear: direct, multi-arm comparisons that test which tools help which learners, under which teaching conditions, and whether the gains persist.

相關論文

大學生向教師解釋幾何作圖,同學在旁思考,桌上的電腦展示相關數碼圖解
政策 / 倫理2026年9月7日
政策 / 倫理 112

評論:Astra 的 AGI 主張,讓教育更需要看見人的真實學習

AIED.HK Editorial

AI Product News Commentary

OpenAI 於 2026 年 9 月 3 日發布 GPT-6 Astra,並引發關於 AGI 是否已經到來的討論。本文將這一說法視為需要歸屬的主張,而非已確立的共識。教育眼前的挑戰,是分清 AI 能產出甚麼,以及學習者能獨立解釋、質疑和遷移甚麼,再以更強的代理能力支援真正的學習。

產品新聞評論GPT-6 Astra
閱讀 500 字摘要 →
三位教育與軟件同事在明亮的大學設計工作室審查圖解教材卡、註釋圖表和數碼原型
政策 / 倫理2026年9月7日
政策 / 倫理 113

評論:Fable 5.1 把更長程的 AI 工作帶進 AIED,教育驗證更顯重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 於 2026 年 9 月 1 日發布 Claude Fable 5.1,強化長程編程與知識工作能力,並降低快取讀取價格。AIED 的機會,是加快從教學構想到可審查原型與研究分析的循環;真正的考驗,是能否把速度轉化為更好的教學與可信證據,同時計入總成本、資料條件和人工審核。

產品新聞評論Claude Fable 5.1
閱讀 500 字摘要 →
A lecturer and two university students inspect ranked learning tools, separate cloud and local plugin cards, and a review ledger in a bright computing studio
政策 / 倫理2026年8月23日
政策 / 倫理 111

Product news: ChatGPT plugin ranking and Claude Code 2.1.239 make tool selection and workspace boundaries inspectable

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: ChatGPT now ranks plugin recommendations partly by continued use after installation and adds more time-aware answers, while Claude Code 2.1.239 distinguishes cloud-synced plugins from local installations and makes a data-residency cost premium visible. Gemini for Education supplies the institutional purpose boundary across teaching, learning and work. Together, the updates make tool selection, context, cost and human review part of AI workflow literacy.

product newsChatGPT pluginsClaude Code 2.1.239
閱讀 500 字摘要 →