← 返回研究新闻
Editorial cover for a randomized comparison of AI support in university programming education
期刊论文同行评审研究20262026年7月19日· 10 min

Less stress, better scores, same learning: The dissociation of performance and learning in AI-supported programming education

Patrick Bassner, Ben Lenk-Ostendorf, Ramona Beinstingel, Tobias Wasner, Stephan Krusche

Computers and Education: Artificial Intelligence

500 字摘要

Editorial cover for a randomized comparison of AI support in university programming education

Bassner, Lenk-Ostendorf, Beinstingel, Wasner, and Krusche examine a central problem in AI-supported education: does better performance while using AI mean that students have learned more? Their open-access study reports a three-arm randomized controlled trial in an introductory programming course at the Technical University of Munich. The design compares two forms of AI assistance with a no-AI control and measures task success, conceptual learning, cognitive load, frustration, and motivation.

The 275 participants completed a 90-minute exercise on concurrency in which they implemented a parallel sum using threading. One group used Iris, a scaffolded tutor designed to provide calibrated hints while withholding complete solutions. A second group used unrestricted ChatGPT, which could return full solutions. The control group used conventional web resources without an AI assistant. This contrast is valuable because it separates the presence of generative AI from the instructional design wrapped around it.

The researchers measured performance through automated test coverage for the programming exercise. Learning was assessed through pre- and post-knowledge tests and a code-comprehension task. Validated scales captured intrinsic, germane, and extraneous cognitive load, frustration, and intrinsic motivation. These separate measures let the study test whether students who complete more of the exercise also show greater conceptual gains, rather than treating a working submission as sufficient evidence of learning.

Both AI-supported groups achieved substantially higher exercise scores than the control group. The score distributions differed: ChatGPT users clustered toward high task scores, control participants clustered lower, and Iris users were spread more widely across the performance range. Students in both AI conditions also reported less frustration and lower extraneous and germane cognitive load than the control group, while intrinsic cognitive load did not differ. Participants rated ChatGPT as easier to use and more helpful.

The learning results tell a more cautious story. Neither AI condition produced larger pre-to-post knowledge gains or a code-comprehension advantage over conventional resources. In this short programming exercise, generative AI therefore functioned mainly as a performance aid. The scaffolded Iris tutor did show one distinctive benefit: it increased intrinsic motivation, while unrestricted ChatGPT did not. The authors describe the attraction of easy, helpful assistance without additional learning as a comfort trap, where learners' preferences and immediate success can diverge from pedagogical effectiveness.

The study should not be generalized beyond its boundaries. It examines one 90-minute concurrency task in one introductory university programming context, and its learning measures focus on knowledge change and code comprehension around that activity. Longer courses, different domains, repeated tutor use, or designs that require explanation and transfer may produce different outcomes. The open dataset and analysis materials strengthen transparency and make replication or secondary analysis possible, but broader evidence is still needed.

For AIEDHK, the practical implication is to separate assistance metrics from learning metrics. A school or university pilot should not infer understanding from completion rate, code quality, reduced frustration, or positive user ratings alone. AI-supported tasks can require students to predict program behavior, explain a generated solution, debug a novel case, compare alternatives, and complete an independent transfer assessment. Scaffolded hints may better preserve motivation and productive activity than unrestricted answer delivery, but their value still needs to be demonstrated through learning evidence. The paper's most important contribution is a clear warning: when AI makes work easier, assessment must work harder to reveal what the learner can actually understand and do.

相关论文

大学生向教师解释几何作图,同学在旁思考,桌上的电脑展示相关数字图解
政策 / 伦理2026年9月7日
政策 / 伦理 112

评论:Astra 的 AGI 主张,让教育更需要看见人的真实学习

AIED.HK Editorial

AI Product News Commentary

OpenAI 于 2026 年 9 月 3 日发布 GPT-6 Astra,并引发关于 AGI 是否已经到来的讨论。本文将这一说法视为需要归属的主张,而非已确立的共识。教育眼前的挑战,是分清 AI 能产出什么,以及学习者能独立解释、质疑和迁移什么,再以更强的代理能力支持真正的学习。

产品新闻评论GPT-6 Astra
阅读 500 字摘要 →
三位教育与软件同事在明亮的大学设计工作室审查图解教材卡、注释图表和数字原型
政策 / 伦理2026年9月7日
政策 / 伦理 113

评论:Fable 5.1 把更长程的 AI 工作带进 AIED,教育验证更显重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 于 2026 年 9 月 1 日发布 Claude Fable 5.1,强化长程编程与知识工作能力,并降低缓存读取价格。AIED 的机会,是加快从教学构想到可审查原型与研究分析的循环;真正的考验,是能否把速度转化为更好的教学与可信证据,同时计入总成本、数据条件和人工审核。

产品新闻评论Claude Fable 5.1
阅读 500 字摘要 →
A lecturer and two university students inspect ranked learning tools, separate cloud and local plugin cards, and a review ledger in a bright computing studio
政策 / 伦理2026年8月23日
政策 / 伦理 111

Product news: ChatGPT plugin ranking and Claude Code 2.1.239 make tool selection and workspace boundaries inspectable

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: ChatGPT now ranks plugin recommendations partly by continued use after installation and adds more time-aware answers, while Claude Code 2.1.239 distinguishes cloud-synced plugins from local installations and makes a data-residency cost premium visible. Gemini for Education supplies the institutional purpose boundary across teaching, learning and work. Together, the updates make tool selection, context, cost and human review part of AI workflow literacy.

product newsChatGPT pluginsClaude Code 2.1.239
阅读 500 字摘要 →