← 返回研究新闻
A diverse group of computing students and an instructor compare AI tutoring responses with a prerequisite map and scaffolded Python exercises
会议论文会议论文20262026年8月10日· 8 min

Pedagogical-fit feedback improved 51 of 62 weak AI-tutor responses, but scaffolding gains created trade-offs

Benjamin Barlog, Hudson Craig, Zedong Peng

IEEE 27th International Conference on Information Reuse and Integration for Data Science (IRI 2026)

500 字摘要

A diverse group of computing students and an instructor compare AI tutoring responses with a prerequisite map and scaffolded Python exercises

Barlog, Craig and Peng ask a sharper question than whether an AI tutor gives a correct answer: is the help appropriate for this learner at this point in a course? Their paper, presented at IEEE IRI 2026 and posted to arXiv on August 5, introduces the Pedagogical Suitability Index, or PSI. The study evaluates ChatGPT, Gemini, Gemma 4 and Qwen 3 in an introductory Python course. Across 30 scenarios, paired standard and deliberately defective student prompts and four models, the authors produced 240 tutor-response evaluations.

PSI combines six equally weighted components: knowledge distance from the learner’s current foundation, prerequisite-order violations, scaffolding density, retention timing, avoidable cognitive load and alignment with the intended Bloom level. The course model included 85 Python concepts and their prerequisites. Scenarios represented two questions for each week of a 15-week course and included learner profiles, common misconceptions and eight prompt-defect categories such as missing context, vague errors and wrong terminology. Two former students rated the scenarios for realism and course fit, although this was a small validation step.

Baseline PSI scores ranged from 0.557 to 0.638. ChatGPT scored highest in this implementation, followed by Qwen 3, Gemma 4 and Gemini, but the spread was modest and did not support a simple closed-versus-open model conclusion. Overall PSI barely changed under defective prompts, declining by 0.002, yet the sub-scores moved in different directions. Missing context reduced knowledge calibration and scaffolding while prerequisite ordering improved, suggesting that a stable composite can hide educationally important trade-offs.

The researchers then selected 62 weak defective-prompt cases for one round of PSI-guided regeneration. The feedback included the original context, the first response, all six scores, a diagnosis and an improvement checklist. Fifty-one of 62 cases improved, or 82.3 percent, and mean PSI rose by 0.049. Scaffolding density produced the largest change, increasing by 0.272 as responses added worked examples, steps and comprehension checks. A single instructor judged 46 of 62 regenerations meaningfully better and 51 of 62 to have resolved the flagged weakness. However, prerequisite-order and Bloom-alignment scores declined on average, so adding visible support did not improve every pedagogical dimension.

The evidence is promising but preliminary. PSI relies on keyword and regular-expression proxies, uses unvalidated equal weights and has two components with little discriminative power in this implementation. It covers one course, 30 scenarios, one model snapshot and no live student learning outcomes; manual review involved one expert and only the selected weak cases. The model names also describe tested product snapshots, not permanent rankings, so later versions should be re-evaluated rather than assumed to preserve the same order. For Hong Kong education, the practical lesson is to evaluate tutor responses against curriculum sequence, learner readiness and intended cognition, then inspect the components rather than trusting one score. A pilot should combine teacher review with authentic student interactions and independent measures of what learners understand and can transfer after the tutor is removed.

相关论文

大学生向教师解释几何作图,同学在旁思考,桌上的电脑展示相关数字图解
政策 / 伦理2026年9月7日
政策 / 伦理 112

评论:Astra 的 AGI 主张,让教育更需要看见人的真实学习

AIED.HK Editorial

AI Product News Commentary

OpenAI 于 2026 年 9 月 3 日发布 GPT-6 Astra,并引发关于 AGI 是否已经到来的讨论。本文将这一说法视为需要归属的主张,而非已确立的共识。教育眼前的挑战,是分清 AI 能产出什么,以及学习者能独立解释、质疑和迁移什么,再以更强的代理能力支持真正的学习。

产品新闻评论GPT-6 Astra
阅读 500 字摘要 →
三位教育与软件同事在明亮的大学设计工作室审查图解教材卡、注释图表和数字原型
政策 / 伦理2026年9月7日
政策 / 伦理 113

评论:Fable 5.1 把更长程的 AI 工作带进 AIED,教育验证更显重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 于 2026 年 9 月 1 日发布 Claude Fable 5.1,强化长程编程与知识工作能力,并降低缓存读取价格。AIED 的机会,是加快从教学构想到可审查原型与研究分析的循环;真正的考验,是能否把速度转化为更好的教学与可信证据,同时计入总成本、数据条件和人工审核。

产品新闻评论Claude Fable 5.1
阅读 500 字摘要 →
A learning scientist and two diverse student researchers compare coded learner actions with a structured reasoning map in a university lab
会议论文2026
会议论文 98

INSIDE aligned simulated student actions with internal dialogue, but reasoning fidelity remained partial

Rose Niousha, Minwoo Kang, Narges Norouzi

Conference on Language Modeling 2026

Niousha, Kang and Norouzi introduce INSIDE, a framework that fine-tunes LLM student simulators on paired internal-dialogue traces and observable actions across cognitive, affective, and action dimensions. It improved action fidelity and achieved reasoning alignment up to 57.9 percent across evaluated models. The result advances simulator evaluation but leaves substantial mismatch and does not justify replacing trials with real learners.

student simulationinternal dialogueBloom's taxonomy
阅读 500 字摘要 →