← 返回研究新聞
Editorial cover for LLM feedback inside an intelligent tutoring system
期刊論文同行評審研究20252026年6月23日· 8 min

Generating In-Context, Personalized Feedback for Intelligent Tutors with Large Language Models

Jennifer M. Reddig, Arav Arora, Christopher J. MacLellan

International Journal of Artificial Intelligence in Education

500 字摘要

Editorial cover for LLM feedback inside an intelligent tutoring system

Reddig, Arora, and MacLellan study a question that sits directly at the current boundary between classic intelligent tutoring systems and generative AI: can a large language model generate useful, in-context feedback for learners inside an ITS? The paper focuses on GPT-4 and the Apprentice Tutor College Algebra ITS. Instead of asking whether a chatbot can generally answer math questions, the authors ground the model in tutor context and student error data. That makes the article especially valuable for AIEDHK because it treats generative AI as a component inside a structured tutoring system rather than as a stand-alone teacher.

The paper examines three linked tasks. First, the LLM needs to diagnose student errors. Second, it needs to generate corrective feedback that responds to the specific error. Third, the system needs a way to assess whether the diagnosis and feedback are accurate and helpful. This workflow matters because feedback generation is one of the most tempting applications of LLMs in education. Teachers and tutor authors spend substantial time crafting hints, explanations, and bug-specific messages. If an LLM could reliably produce such feedback, it might reduce authoring burden and make tutoring systems more responsive to unusual learner responses.

The results are deliberately mixed. The study reports that GPT-4 can diagnose a range of student errors, but performance drops when responses are more complex or contain multiple problems. It also finds that generated feedback is often relevant and specific, yet a substantial share of hints are too general, incorrect, or reveal the answer. The authors also test whether an LLM can help evaluate generated feedback. That automated quality-control path is promising, but the reported helpfulness pass rate is low enough to show that autonomous feedback pipelines are not ready to be trusted without stronger review.

This makes the paper useful precisely because it resists a simple pro-AI or anti-AI conclusion. It shows that LLMs can add flexibility to ITS feedback, especially when they are given structured context from the tutor. At the same time, it documents the risk of misleading hints, overhelping, and weak automated evaluation. For research translation, the key point is that feedback quality is a safety issue. A fluent hint can still harm learning if it points to the wrong rule, masks a misconception, or gives away the solution before the learner has done the reasoning.

For AIEDHK, the article can become a practical checklist for LLM tutor design. A trustworthy system should specify the tutor context supplied to the model, the student error categories it can diagnose, the threshold for showing feedback, the human or automated review process, and the policy for withholding low-confidence hints. The paper also supports a hybrid design direction for Hong Kong products: combine the curriculum structure and learner modeling of ITS with the language flexibility of LLMs, while keeping teacher oversight and evidence-based quality checks visible. The strongest message is not that LLMs replace tutor authoring, but that they may extend it when embedded in a disciplined tutoring architecture.

相關論文

大學生向教師解釋幾何作圖,同學在旁思考,桌上的電腦展示相關數碼圖解
政策 / 倫理2026年9月7日
政策 / 倫理 112

評論:Astra 的 AGI 主張,讓教育更需要看見人的真實學習

AIED.HK Editorial

AI Product News Commentary

OpenAI 於 2026 年 9 月 3 日發布 GPT-6 Astra,並引發關於 AGI 是否已經到來的討論。本文將這一說法視為需要歸屬的主張,而非已確立的共識。教育眼前的挑戰,是分清 AI 能產出甚麼,以及學習者能獨立解釋、質疑和遷移甚麼,再以更強的代理能力支援真正的學習。

產品新聞評論GPT-6 Astra
閱讀 500 字摘要 →
三位教育與軟件同事在明亮的大學設計工作室審查圖解教材卡、註釋圖表和數碼原型
政策 / 倫理2026年9月7日
政策 / 倫理 113

評論:Fable 5.1 把更長程的 AI 工作帶進 AIED,教育驗證更顯重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 於 2026 年 9 月 1 日發布 Claude Fable 5.1,強化長程編程與知識工作能力,並降低快取讀取價格。AIED 的機會,是加快從教學構想到可審查原型與研究分析的循環;真正的考驗,是能否把速度轉化為更好的教學與可信證據,同時計入總成本、資料條件和人工審核。

產品新聞評論Claude Fable 5.1
閱讀 500 字摘要 →
Academic cover for a position paper on ChatGPT and large language models in education
綜述2023
綜述 c3c9681c-ca00-4df2-aa39-490f775df4fb

ChatGPT for good? On opportunities and challenges of large language models for education

Enkelejda Kasneci, Kathrin Sessler, Stefan Kuechemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Guennemann, Eyke Huellermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Juergen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuehn, Gjergji Kasneci

Learning and Individual Differences

A widely cited position paper that balances the educational opportunities of large language models with risks around bias, privacy, assessment, and teacher guidance.

large language modelsChatGPTteacher support
閱讀 500 字摘要 →