← 返回研究新闻
Editorial cover for LLM feedback inside an intelligent tutoring system
期刊论文同行评审研究20252026年6月23日· 8 min

Generating In-Context, Personalized Feedback for Intelligent Tutors with Large Language Models

Jennifer M. Reddig, Arav Arora, Christopher J. MacLellan

International Journal of Artificial Intelligence in Education

500 字摘要

Editorial cover for LLM feedback inside an intelligent tutoring system

Reddig, Arora, and MacLellan study a question that sits directly at the current boundary between classic intelligent tutoring systems and generative AI: can a large language model generate useful, in-context feedback for learners inside an ITS? The paper focuses on GPT-4 and the Apprentice Tutor College Algebra ITS. Instead of asking whether a chatbot can generally answer math questions, the authors ground the model in tutor context and student error data. That makes the article especially valuable for AIEDHK because it treats generative AI as a component inside a structured tutoring system rather than as a stand-alone teacher.

The paper examines three linked tasks. First, the LLM needs to diagnose student errors. Second, it needs to generate corrective feedback that responds to the specific error. Third, the system needs a way to assess whether the diagnosis and feedback are accurate and helpful. This workflow matters because feedback generation is one of the most tempting applications of LLMs in education. Teachers and tutor authors spend substantial time crafting hints, explanations, and bug-specific messages. If an LLM could reliably produce such feedback, it might reduce authoring burden and make tutoring systems more responsive to unusual learner responses.

The results are deliberately mixed. The study reports that GPT-4 can diagnose a range of student errors, but performance drops when responses are more complex or contain multiple problems. It also finds that generated feedback is often relevant and specific, yet a substantial share of hints are too general, incorrect, or reveal the answer. The authors also test whether an LLM can help evaluate generated feedback. That automated quality-control path is promising, but the reported helpfulness pass rate is low enough to show that autonomous feedback pipelines are not ready to be trusted without stronger review.

This makes the paper useful precisely because it resists a simple pro-AI or anti-AI conclusion. It shows that LLMs can add flexibility to ITS feedback, especially when they are given structured context from the tutor. At the same time, it documents the risk of misleading hints, overhelping, and weak automated evaluation. For research translation, the key point is that feedback quality is a safety issue. A fluent hint can still harm learning if it points to the wrong rule, masks a misconception, or gives away the solution before the learner has done the reasoning.

For AIEDHK, the article can become a practical checklist for LLM tutor design. A trustworthy system should specify the tutor context supplied to the model, the student error categories it can diagnose, the threshold for showing feedback, the human or automated review process, and the policy for withholding low-confidence hints. The paper also supports a hybrid design direction for Hong Kong products: combine the curriculum structure and learner modeling of ITS with the language flexibility of LLMs, while keeping teacher oversight and evidence-based quality checks visible. The strongest message is not that LLMs replace tutor authoring, but that they may extend it when embedded in a disciplined tutoring architecture.

相关论文

大学生向教师解释几何作图,同学在旁思考,桌上的电脑展示相关数字图解
政策 / 伦理2026年9月7日
政策 / 伦理 112

评论:Astra 的 AGI 主张,让教育更需要看见人的真实学习

AIED.HK Editorial

AI Product News Commentary

OpenAI 于 2026 年 9 月 3 日发布 GPT-6 Astra,并引发关于 AGI 是否已经到来的讨论。本文将这一说法视为需要归属的主张,而非已确立的共识。教育眼前的挑战,是分清 AI 能产出什么,以及学习者能独立解释、质疑和迁移什么,再以更强的代理能力支持真正的学习。

产品新闻评论GPT-6 Astra
阅读 500 字摘要 →
三位教育与软件同事在明亮的大学设计工作室审查图解教材卡、注释图表和数字原型
政策 / 伦理2026年9月7日
政策 / 伦理 113

评论:Fable 5.1 把更长程的 AI 工作带进 AIED,教育验证更显重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 于 2026 年 9 月 1 日发布 Claude Fable 5.1,强化长程编程与知识工作能力,并降低缓存读取价格。AIED 的机会,是加快从教学构想到可审查原型与研究分析的循环;真正的考验,是能否把速度转化为更好的教学与可信证据,同时计入总成本、数据条件和人工审核。

产品新闻评论Claude Fable 5.1
阅读 500 字摘要 →
Academic cover for a position paper on ChatGPT and large language models in education
综述2023
综述 c3c9681c-ca00-4df2-aa39-490f775df4fb

ChatGPT for good? On opportunities and challenges of large language models for education

Enkelejda Kasneci, Kathrin Sessler, Stefan Kuechemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Guennemann, Eyke Huellermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Juergen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuehn, Gjergji Kasneci

Learning and Individual Differences

A widely cited position paper that balances the educational opportunities of large language models with risks around bias, privacy, assessment, and teacher guidance.

large language modelsChatGPTteacher support
阅读 500 字摘要 →