
GPT-4 feedback increases student activation and learning outcomes in higher education
Stephan Geschwind, Johann Graf Lambsdorff, Deborah Voss, Veronika Hackl
International Journal of Artificial Intelligence in Education
500語要約

Geschwind, Lambsdorff, Voss, and Hackl investigate whether generative AI can make individualized feedback more scalable in large university courses. Their 2026 open-access IJAIED article reports a lab-in-the-field experiment conducted in undergraduate macroeconomics tutorial classes over one semester. Students answered eight open-ended questions and experienced one of three feedback arrangements: classroom-level lecturer feedback, additional individual feedback from peers, or additional individual feedback generated by GPT-4. The study asks whether the feedback changes student activation and the quality of subsequent answers, rather than only whether students say they like AI.
All groups received lecturer discussion of a sample solution and adaptive feedback on two selected student answers. In the peer-feedback condition, students anonymously rated another student's answer for content and style and wrote suggestions for improvement. In the AI-feedback condition, GPT-4 provided individual scores and qualitative suggestions in a comparable structure. Both individual-feedback formats were designed around three familiar feedback functions: looking back at current performance, clarifying the goal through a sample solution, and pointing forward to improvement.
The researchers operationalized activation in two ways. First, they tracked voluntary participation across the eight tasks; tutorial participation did not count toward the final examination, although students who completed at least seven tasks entered a raffle. Second, they used answer length as an indicator of the effort invested in a task. For learning outcomes, GPT-4 rated the content and style of answers on five-point scales. To make treatment comparisons more consistent, the model rated answers from all three conditions after the course in three separate iterations, and the researchers averaged those ratings.
The AI-feedback condition produced the clearest activation pattern. Participation remained higher over time than in the lecturer-only condition, while peer feedback performed only marginally better than lecturer feedback. Students receiving AI feedback also wrote the longest answers. Because attrition was voluntary and not random, the authors did not rely only on simple averages: their task-to-task analysis used 479 repeated comparisons in which the same student participated in the same condition on consecutive relevant tasks.
For learning outcomes, AI feedback was associated with the strongest improvement in content ratings. Peer feedback did not produce the same advantage, which the authors connect to differences in reliability and completeness: peers sometimes failed to provide feedback, while the AI returned it consistently. The study found no comparable treatment effect on writing style. That distinction matters because it suggests that the intervention supported substantive engagement with macroeconomic answers without automatically improving every dimension of academic writing.
The paper is promising but should be interpreted carefully. The setting was one subject area and one university course, participation was voluntary, attrition differed across conditions, and the outcome measure was improvement in open-ended answers rather than a broad examination of long-term retention or transfer. GPT-4 generated feedback in one condition and also served as the common post-course rater, even though the authors used repeated ratings and prior reliability work to strengthen the measure. These choices make the study more informative than a satisfaction survey, but they do not establish that AI feedback will outperform expert human feedback in every context.
For AIEDHK, the practical lesson is to focus on feedback system design. Individual AI feedback may sustain participation when it is timely, structured, connected to a sample solution, and followed by opportunities to revise. Educators should preserve lecturer oversight, test feedback validity, measure learning with independent assessments, and compare the intervention with realistic alternatives. The contribution is not a claim that GPT-4 replaces teachers. It is evidence that carefully structured AI feedback can extend the reach of tutorial support while keeping course goals and human judgment visible.


