← 返回研究新聞
Editorial cover for a review of generative AI in programming education
綜述證據綜述20252026年6月25日· 8 min

Literature Review on the Integration of Generative AI in Programming Education

Jemimah Nathaniel, Solomon Sunday Oyelere, Jarkko Suhonen, Matti Tedre

International Journal of Artificial Intelligence in Education

500 字摘要

Editorial cover for a review of generative AI in programming education

Nathaniel, Oyelere, Suhonen, and Tedre review a question that is now central to computer science education: how can generative AI tools be integrated into programming education without weakening students' foundational logic, problem solving, and higher-order thinking skills? The paper is useful for AIEDHK because it moves beyond generic enthusiasm for ChatGPT or Copilot. It asks whether the tools are embedded in teaching methods, assessment routines, and learning processes that still require students to understand code rather than only generate it.

The review synthesizes 40 empirical studies using PRISMA 2020 and Kitchenham-style review methods. Its focus is not simply whether GenAI can solve programming tasks. Instead, it examines how studies connect GenAI tools with programming curricula, teaching methods, assessment designs, integration processes, and student skill development. That framing is important because programming education has a long history of intelligent tutoring systems, automated feedback, Parsons problems, code explanation tools, and step-based support. GenAI adds flexibility and natural-language interaction, but it also increases the risk that learners accept generated code without understanding algorithms, syntax, data structures, or debugging logic.

The paper's findings are deliberately implementation-focused. The authors argue that successful integration depends on intentional teaching strategies, thoughtfully designed assessments, and structured integration processes. They also identify barriers: limited accessibility support, insufficient bias mitigation, weak curriculum alignment, and tool selection driven by availability rather than educational fit. These are practical concerns for any school or university considering AI-assisted coding. A tool that improves productivity for experienced developers can still be harmful for novice learners if it bypasses the struggle needed to build mental models.

The review also proposes a GenAI-Ped framework that combines self-regulated learning, universal design principles, prompt-engineering support, and iterative feedback. For AIEDHK, this is the most actionable contribution. It suggests that GenAI in coding courses should be framed as a guided learning partner, not an answer machine. Students need orientation on when to ask for help, how to inspect generated code, how to explain a solution, and how to reflect on what they have learned. Teachers need assessment formats that reveal reasoning, not only final code output. Product teams need interfaces that encourage explanation, comparison, revision, and metacognitive checks.

The paper is especially relevant for Hong Kong because programming education is multilingual, high-stakes, and often linked to future workforce claims. GenAI coding support can make programming more accessible, but only if it is aligned with local curricula, language needs, teacher capacity, and assessment expectations. AIEDHK can use this review to evaluate AI coding tutors, coding assistants, and student copilots through a clear test: does the system help learners develop durable programming understanding, or does it mainly make correct-looking code easier to obtain?

相關論文

大學生向教師解釋幾何作圖,同學在旁思考,桌上的電腦展示相關數碼圖解
政策 / 倫理2026年9月7日
政策 / 倫理 112

評論:Astra 的 AGI 主張,讓教育更需要看見人的真實學習

AIED.HK Editorial

AI Product News Commentary

OpenAI 於 2026 年 9 月 3 日發布 GPT-6 Astra,並引發關於 AGI 是否已經到來的討論。本文將這一說法視為需要歸屬的主張,而非已確立的共識。教育眼前的挑戰,是分清 AI 能產出甚麼,以及學習者能獨立解釋、質疑和遷移甚麼,再以更強的代理能力支援真正的學習。

產品新聞評論GPT-6 Astra
閱讀 500 字摘要 →
A programming lecturer and two diverse university students inspect compiled code, an inheritance diagram, and a grading rubric in a computer laboratory
期刊論文2026
期刊論文 102

Five AI systems outscored the average OOP cohort but still failed compilation and advanced concepts

Marina Lepp, Joosep Kaimre

arXiv preprint

Lepp and Kaimre evaluated ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and Microsoft 365 Copilot on authentic introductory OOP tests and examinations using student grading criteria. Systems exceeded the historical average and often solved long tasks, yet some code did not compile and interfaces, abstract classes, inheritance, and image-based questions remained difficult. The results challenge take-home assessment validity without proving student learning.

programming assessmentobject-oriented programminggenerative AI
閱讀 500 字摘要 →
A Black university student debugs from her own notes while an instructor supports her beside a graduated cyan help ladder whose final solution rung is locked
工具 / 數據集2026
工具 / 數據集 92

A guarded LLM tutor reached its withholding targets in scripted tests, but student learning remains unmeasured

Yusuf Pisan

arXiv preprint

Pisan reports a deployed programming tutor that places an eight-rung help ceiling outside the generating LLM, strips solution code deterministically and judges risky replies against a per-turn contract. Across roughly two dozen scripted turns per calibration run, earnest-reply revisions fell from 43% to 0% and audited ceiling compliance rose from about 77% to 100% after measurement and policy defects were repaired. The evaluation used synthetic personas, not students, so it establishes contract compliance rather than durable learning.

Socratic tutoringanswer withholdingprogramming education
閱讀 500 字摘要 →