← গবেষণা সংবাদে ফিরুন
Editorial cover for a randomized comparison of AI support in university programming education
জার্নাল পেপারPeer-reviewed study2026১৯ জুল, ২০২৬· 3 min

Less stress, better scores, same learning: The dissociation of performance and learning in AI-supported programming education

Patrick Bassner, Ben Lenk-Ostendorf, Ramona Beinstingel, Tobias Wasner, Stephan Krusche

Computers and Education: Artificial Intelligence

৫০০-শব্দের সারাংশ

Editorial cover for a randomized comparison of AI support in university programming education

Bassner, Lenk-Ostendorf, Beinstingel, Wasner, and Krusche examine a central problem in AI-supported education: does better performance while using AI mean that students have learned more? Their open-access study reports a three-arm randomized controlled trial in an introductory programming course at the Technical University of Munich. The design compares two forms of AI assistance with a no-AI control and measures task success, conceptual learning, cognitive load, frustration, and motivation.

The 275 participants completed a 90-minute exercise on concurrency in which they implemented a parallel sum using threading. One group used Iris, a scaffolded tutor designed to provide calibrated hints while withholding complete solutions. A second group used unrestricted ChatGPT, which could return full solutions. The control group used conventional web resources without an AI assistant. This contrast is valuable because it separates the presence of generative AI from the instructional design wrapped around it.

The researchers measured performance through automated test coverage for the programming exercise. Learning was assessed through pre- and post-knowledge tests and a code-comprehension task. Validated scales captured intrinsic, germane, and extraneous cognitive load, frustration, and intrinsic motivation. These separate measures let the study test whether students who complete more of the exercise also show greater conceptual gains, rather than treating a working submission as sufficient evidence of learning.

Both AI-supported groups achieved substantially higher exercise scores than the control group. The score distributions differed: ChatGPT users clustered toward high task scores, control participants clustered lower, and Iris users were spread more widely across the performance range. Students in both AI conditions also reported less frustration and lower extraneous and germane cognitive load than the control group, while intrinsic cognitive load did not differ. Participants rated ChatGPT as easier to use and more helpful.

The learning results tell a more cautious story. Neither AI condition produced larger pre-to-post knowledge gains or a code-comprehension advantage over conventional resources. In this short programming exercise, generative AI therefore functioned mainly as a performance aid. The scaffolded Iris tutor did show one distinctive benefit: it increased intrinsic motivation, while unrestricted ChatGPT did not. The authors describe the attraction of easy, helpful assistance without additional learning as a comfort trap, where learners' preferences and immediate success can diverge from pedagogical effectiveness.

The study should not be generalized beyond its boundaries. It examines one 90-minute concurrency task in one introductory university programming context, and its learning measures focus on knowledge change and code comprehension around that activity. Longer courses, different domains, repeated tutor use, or designs that require explanation and transfer may produce different outcomes. The open dataset and analysis materials strengthen transparency and make replication or secondary analysis possible, but broader evidence is still needed.

For AIEDHK, the practical implication is to separate assistance metrics from learning metrics. A school or university pilot should not infer understanding from completion rate, code quality, reduced frustration, or positive user ratings alone. AI-supported tasks can require students to predict program behavior, explain a generated solution, debug a novel case, compare alternatives, and complete an independent transfer assessment. Scaffolded hints may better preserve motivation and productive activity than unrestricted answer delivery, but their value still needs to be demonstrated through learning evidence. The paper's most important contribution is a clear warning: when AI makes work easier, assessment must work harder to reveal what the learner can actually understand and do.

সম্পর্কিত পেপার

A university student explains a geometry construction to a lecturer while a classmate follows and a laptop displays a related digital diagram
নীতি / নৈতিকতা৭ সেপ, ২০২৬
নীতি / নৈতিকতা 112

Commentary: Astra's AGI claim puts evidence of human learning at the centre of education

AIED.HK Editorial

AI Product News Commentary

OpenAI launched GPT-6 Astra on 3 September 2026 amid claims about the arrival of AGI. This commentary treats that label as a claim, not an established consensus. For education, the immediate challenge is to distinguish what an AI can produce from what a learner can explain, question and transfer independently—and to use stronger agents to support that learning.

product newscommentaryGPT-6 Astra
৫০০-শব্দের সারাংশ পড়ুন →
Three education and software colleagues review illustrated lesson cards, an annotated chart and a digital prototype in a bright university design studio
নীতি / নৈতিকতা৭ সেপ, ২০২৬
নীতি / নৈতিকতা 113

Commentary: Fable 5.1 brings longer AI workflows to AIED—and makes educational validation more important

AIED.HK Editorial

AI Product News Commentary

Anthropic released Claude Fable 5.1 on 1 September 2026 with stronger long-running coding and knowledge-work capabilities and cheaper cache reads. For AIED, the opportunity is a faster cycle from teaching idea to reviewable prototype and research analysis. The test is whether teams can turn that speed into better pedagogy and credible evidence, while accounting for total cost, data conditions and human review.

product newscommentaryClaude Fable 5.1
৫০০-শব্দের সারাংশ পড়ুন →
A lecturer and two university students inspect ranked learning tools, separate cloud and local plugin cards, and a review ledger in a bright computing studio
নীতি / নৈতিকতা২৩ আগ, ২০২৬
নীতি / নৈতিকতা 111

Product news: ChatGPT plugin ranking and Claude Code 2.1.239 make tool selection and workspace boundaries inspectable

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: ChatGPT now ranks plugin recommendations partly by continued use after installation and adds more time-aware answers, while Claude Code 2.1.239 distinguishes cloud-synced plugins from local installations and makes a data-residency cost premium visible. Gemini for Education supplies the institutional purpose boundary across teaching, learning and work. Together, the updates make tool selection, context, cost and human review part of AI workflow literacy.

product newsChatGPT pluginsClaude Code 2.1.239
৫০০-শব্দের সারাংশ পড়ুন →