গবেষণা সংবাদে ফিরুন
Editorial cover for LLM feedback inside an intelligent tutoring system
জার্নাল পেপারPeer-reviewed study2025২৩ জুন, ২০২৬· 2 min

Generating In-Context, Personalized Feedback for Intelligent Tutors with Large Language Models

Jennifer M. Reddig, Arav Arora, Christopher J. MacLellan

International Journal of Artificial Intelligence in Education

৫০০-শব্দের সারাংশ

Premium tabletop illustration of LLM-generated tutoring feedback passing through error diagnosis, helpfulness review, confidence gates, and teacher oversight.

Reddig, Arora, and MacLellan study a question that sits directly at the current boundary between classic intelligent tutoring systems and generative AI: can a large language model generate useful, in-context feedback for learners inside an ITS? The paper focuses on GPT-4 and the Apprentice Tutor College Algebra ITS. Instead of asking whether a chatbot can generally answer math questions, the authors ground the model in tutor context and student error data. That makes the article especially valuable for AIEDHK because it treats generative AI as a component inside a structured tutoring system rather than as a stand-alone teacher.

The paper examines three linked tasks. First, the LLM needs to diagnose student errors. Second, it needs to generate corrective feedback that responds to the specific error. Third, the system needs a way to assess whether the diagnosis and feedback are accurate and helpful. This workflow matters because feedback generation is one of the most tempting applications of LLMs in education. Teachers and tutor authors spend substantial time crafting hints, explanations, and bug-specific messages. If an LLM could reliably produce such feedback, it might reduce authoring burden and make tutoring systems more responsive to unusual learner responses.

The results are deliberately mixed. The study reports that GPT-4 can diagnose a range of student errors, but performance drops when responses are more complex or contain multiple problems. It also finds that generated feedback is often relevant and specific, yet a substantial share of hints are too general, incorrect, or reveal the answer. The authors also test whether an LLM can help evaluate generated feedback. That automated quality-control path is promising, but the reported helpfulness pass rate is low enough to show that autonomous feedback pipelines are not ready to be trusted without stronger review.

This makes the paper useful precisely because it resists a simple pro-AI or anti-AI conclusion. It shows that LLMs can add flexibility to ITS feedback, especially when they are given structured context from the tutor. At the same time, it documents the risk of misleading hints, overhelping, and weak automated evaluation. For research translation, the key point is that feedback quality is a safety issue. A fluent hint can still harm learning if it points to the wrong rule, masks a misconception, or gives away the solution before the learner has done the reasoning.

For AIEDHK, the article can become a practical checklist for LLM tutor design. A trustworthy system should specify the tutor context supplied to the model, the student error categories it can diagnose, the threshold for showing feedback, the human or automated review process, and the policy for withholding low-confidence hints. The paper also supports a hybrid design direction for Hong Kong products: combine the curriculum structure and learner modeling of ITS with the language flexibility of LLMs, while keeping teacher oversight and evidence-based quality checks visible. The strongest message is not that LLMs replace tutor authoring, but that they may extend it when embedded in a disciplined tutoring architecture.

সম্পর্কিত পেপার

Eight diverse adult educators work in three small groups with generic text-free laptops while a mentor guides a practical workshop discussion
নীতি / নৈতিকতা৩০ জুল, ২০২৬
নীতি / নৈতিকতা 49

OpenAI product news: AI Skills Jam brings hands-on AI practice to K-12 educators

OpenAI

AI Product and Learning Report

Product news: OpenAI Academy and the Walton Family Foundation announced hands-on AI Skills Jam workshops for more than 1,600 US K-12 educators and leaders, linking practical experimentation with continuing resources while leaving learning impact to be independently evaluated.

product newsteacher professional learningAI Skills Jam
৫০০-শব্দের সারাংশ পড়ুন
A pre-service chemistry teacher and instructor review a text-free AI-assisted lesson design while a separate intact class works with paper models behind glass
জার্নাল পেপার2026
জার্নাল পেপার 48

Unscaffolded GenAI use in teacher education showed no instructional-design advantage

Jun Zhang, Yuting Peng, Xinyue Deng, Qin Zeng, Kai Wang

Behavioral Sciences

A 2026 quasi-experiment with 52 pre-service chemistry teachers found no adjusted advantage from permitted but unscaffolded GenAI use in AI readiness, self-regulated learning, or critical thinking, while the no-GenAI group achieved stronger instructional-design performance.

pre-service teachersunscaffolded GenAIinstructional design
৫০০-শব্দের সারাংশ পড়ুন
Academic cover for a position paper on ChatGPT and large language models in education
রিভিউ2023
রিভিউ c3c9681c-ca00-4df2-aa39-490f775df4fb

ChatGPT for good? On opportunities and challenges of large language models for education

Enkelejda Kasneci, Kathrin Sessler, Stefan Kuechemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Guennemann, Eyke Huellermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Juergen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuehn, Gjergji Kasneci

Learning and Individual Differences

A widely cited position paper that balances the educational opportunities of large language models with risks around bias, privacy, assessment, and teacher guidance.

large language modelsChatGPTteacher support
৫০০-শব্দের সারাংশ পড়ুন