
ChatGPT for good? On opportunities and challenges of large language models for education
Enkelejda Kasneci, Kathrin Sessler, Stefan Kuechemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Guennemann, Eyke Huellermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Juergen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuehn, Gjergji Kasneci
Learning and Individual Differences
500-word summary

Listen to the 500-word summary
Audio summary
This position paper became highly cited almost immediately because it gave education researchers and practitioners an early, structured way to discuss large language models after the public arrival of ChatGPT. Rather than treating ChatGPT only as a novelty or threat, the authors frame large language models as a durable class of AI systems that will affect student learning, teacher work, assessment, professional training, accessibility, and institutional governance. The paper is valuable because it balances opportunity and caution without collapsing into either hype or prohibition.
The opportunity side of the paper focuses on language-intensive educational work. Large language models can generate explanations, examples, summaries, questions, prompts, feedback, and drafts. They can support reading, writing, mathematics, science, language learning, programming, and professional communication. From the student perspective, the paper highlights personalized practice, adaptive explanations, accessible reformulations, and support for learners with diverse needs. From the teacher perspective, it points to lesson preparation, assessment item generation, feedback drafting, and professional development resources. For AIEDHK, these categories map neatly onto product opportunities in multilingual feedback, teacher workflow support, and learner-facing explanation tools.
The challenge side is equally important. The paper warns that LLM outputs may be biased, inaccurate, non-transparent, privacy-sensitive, or difficult to distinguish from student work. It also notes that educational use cannot be judged only by whether the model produces fluent text. Good deployment requires educator guidance, user interface design, fairness measures, bias correction, transparency mechanisms, training resources, data privacy routines, and regular review. In other words, the model is not the educational intervention by itself; the intervention includes pedagogy, policy, interfaces, review processes, and institutional norms.
For review purposes, the paper is especially helpful because it treats LLM adoption as a systems problem. It asks how teachers, learners, institutions, model developers, and policymakers interact when a general-purpose language model enters educational practice. This matters for AIEDHK because many product ideas sound persuasive at the feature level: generate feedback, rewrite text, create exercises, answer questions. Kasneci and colleagues push the analysis one level deeper. What data are being used? What kind of task is appropriate? Who reviews the output? How are students taught to use the tool critically? What happens to assessment when assistance becomes easy to access? These questions turn a broad LLM discussion into an implementation checklist.
This paper connects generative AI to AIED without pretending that generative AI replaces earlier AIED research. It extends classic concerns about personalization, feedback, assessment, and teacher support into the LLM era. It also gives AIEDHK a bridge from foundational intelligent tutoring research to current public debates about ChatGPT in schools and universities. The central takeaway is that LLMs should be understood as configurable infrastructure for educational tasks, not as autonomous teachers. That framing is useful for school leaders because it makes adoption a question of workflow, policy, and evidence, not just access to a powerful model.

