← 返回研究新聞
A diverse group of computing students and an instructor compare AI tutoring responses with a prerequisite map and scaffolded Python exercises
會議論文會議論文20262026年8月10日· 8 min

Pedagogical-fit feedback improved 51 of 62 weak AI-tutor responses, but scaffolding gains created trade-offs

Benjamin Barlog, Hudson Craig, Zedong Peng

IEEE 27th International Conference on Information Reuse and Integration for Data Science (IRI 2026)

500 字摘要

A diverse group of computing students and an instructor compare AI tutoring responses with a prerequisite map and scaffolded Python exercises

Barlog, Craig and Peng ask a sharper question than whether an AI tutor gives a correct answer: is the help appropriate for this learner at this point in a course? Their paper, presented at IEEE IRI 2026 and posted to arXiv on August 5, introduces the Pedagogical Suitability Index, or PSI. The study evaluates ChatGPT, Gemini, Gemma 4 and Qwen 3 in an introductory Python course. Across 30 scenarios, paired standard and deliberately defective student prompts and four models, the authors produced 240 tutor-response evaluations.

PSI combines six equally weighted components: knowledge distance from the learner’s current foundation, prerequisite-order violations, scaffolding density, retention timing, avoidable cognitive load and alignment with the intended Bloom level. The course model included 85 Python concepts and their prerequisites. Scenarios represented two questions for each week of a 15-week course and included learner profiles, common misconceptions and eight prompt-defect categories such as missing context, vague errors and wrong terminology. Two former students rated the scenarios for realism and course fit, although this was a small validation step.

Baseline PSI scores ranged from 0.557 to 0.638. ChatGPT scored highest in this implementation, followed by Qwen 3, Gemma 4 and Gemini, but the spread was modest and did not support a simple closed-versus-open model conclusion. Overall PSI barely changed under defective prompts, declining by 0.002, yet the sub-scores moved in different directions. Missing context reduced knowledge calibration and scaffolding while prerequisite ordering improved, suggesting that a stable composite can hide educationally important trade-offs.

The researchers then selected 62 weak defective-prompt cases for one round of PSI-guided regeneration. The feedback included the original context, the first response, all six scores, a diagnosis and an improvement checklist. Fifty-one of 62 cases improved, or 82.3 percent, and mean PSI rose by 0.049. Scaffolding density produced the largest change, increasing by 0.272 as responses added worked examples, steps and comprehension checks. A single instructor judged 46 of 62 regenerations meaningfully better and 51 of 62 to have resolved the flagged weakness. However, prerequisite-order and Bloom-alignment scores declined on average, so adding visible support did not improve every pedagogical dimension.

The evidence is promising but preliminary. PSI relies on keyword and regular-expression proxies, uses unvalidated equal weights and has two components with little discriminative power in this implementation. It covers one course, 30 scenarios, one model snapshot and no live student learning outcomes; manual review involved one expert and only the selected weak cases. The model names also describe tested product snapshots, not permanent rankings, so later versions should be re-evaluated rather than assumed to preserve the same order. For Hong Kong education, the practical lesson is to evaluate tutor responses against curriculum sequence, learner readiness and intended cognition, then inspect the components rather than trusting one score. A pilot should combine teacher review with authentic student interactions and independent measures of what learners understand and can transfer after the tutor is removed.

相關論文

大學生向教師解釋幾何作圖,同學在旁思考,桌上的電腦展示相關數碼圖解
政策 / 倫理2026年9月7日
政策 / 倫理 112

評論:Astra 的 AGI 主張,讓教育更需要看見人的真實學習

AIED.HK Editorial

AI Product News Commentary

OpenAI 於 2026 年 9 月 3 日發布 GPT-6 Astra,並引發關於 AGI 是否已經到來的討論。本文將這一說法視為需要歸屬的主張,而非已確立的共識。教育眼前的挑戰,是分清 AI 能產出甚麼,以及學習者能獨立解釋、質疑和遷移甚麼,再以更強的代理能力支援真正的學習。

產品新聞評論GPT-6 Astra
閱讀 500 字摘要 →
三位教育與軟件同事在明亮的大學設計工作室審查圖解教材卡、註釋圖表和數碼原型
政策 / 倫理2026年9月7日
政策 / 倫理 113

評論:Fable 5.1 把更長程的 AI 工作帶進 AIED,教育驗證更顯重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 於 2026 年 9 月 1 日發布 Claude Fable 5.1,強化長程編程與知識工作能力,並降低快取讀取價格。AIED 的機會,是加快從教學構想到可審查原型與研究分析的循環;真正的考驗,是能否把速度轉化為更好的教學與可信證據,同時計入總成本、資料條件和人工審核。

產品新聞評論Claude Fable 5.1
閱讀 500 字摘要 →
A learning scientist and two diverse student researchers compare coded learner actions with a structured reasoning map in a university lab
會議論文2026
會議論文 98

INSIDE aligned simulated student actions with internal dialogue, but reasoning fidelity remained partial

Rose Niousha, Minwoo Kang, Narges Norouzi

Conference on Language Modeling 2026

Niousha, Kang and Norouzi introduce INSIDE, a framework that fine-tunes LLM student simulators on paired internal-dialogue traces and observable actions across cognitive, affective, and action dimensions. It improved action fidelity and achieved reasoning alignment up to 57.9 percent across evaluated models. The result advances simulator evaluation but leaves substantial mismatch and does not justify replacing trials with real learners.

student simulationinternal dialogueBloom's taxonomy
閱讀 500 字摘要 →