
Enhancing School Students' Self-Regulated Learning through Generative AI Support: A Randomized Controlled Trial
Tim Fütterer, Lisa Bardach, Jochen Kuhn, Stefan Daniel Keller, Peter Gerjets
Educational Psychology Review
Resumo de 500 palavras

Fütterer, Bardach, Kuhn, Keller, and Gerjets test whether theory-informed generative-AI prompts can strengthen self-regulated learning in authentic secondary-school lessons. Their open-access article in Educational Psychology Review is valuable because it uses a preregistered randomized design, runs inside regular physics and English classes, and compares two pedagogical interventions with a strong control condition that still had access to standard ChatGPT. The study therefore asks whether carefully targeted prompting adds value beyond ordinary chatbot use, not whether any AI access is better than none.
The sample included 371 students in Grades 7 to 9 from secondary schools in Baden-Württemberg, Germany. The students' mean age was 13.92 years, 54 percent were female, and 45 percent reported a migration background. They were randomly assigned at the individual level to one of three conditions: GPT-supported reflection on the personal utility of the learning content, GPT-supported prompting to use an elaboration strategy through learning by explaining, or a standard ChatGPT control without the pedagogical prompts. The systems all used GPT-4o, so the experimental difference came from the instructional design rather than the underlying model.
Students worked through six 45-minute sessions during regular lessons in April and May 2025. The platform presented subject tasks alongside the chatbot. In the utility-value condition, the system connected content to students' everyday lives, interests, and aspirations. In the cognitive-strategy condition, it encouraged students to explain concepts and elaborate their understanding. The control offered a general conversational assistant. Teachers provided technical support and supervised implementation but did not normally give content help. The prompts, preregistration, data, supplementary materials, and reproducible analysis code are publicly available through the study's OSF materials.
The researchers measured perceived utility value, self-reported strategy use, tested performance on learning by explaining, maintained interest, effort, and domain-specific knowledge before and after the intervention. The main result was cautious. Utility value developed more favourably in the utility-reflection condition than in the cognitive-strategy condition. Yet the utility condition did not differ significantly from standard ChatGPT, and the pattern mainly reflected a decline in the cognitive-strategy group rather than a clear increase in utility value. There were no significant advantages of either pedagogical intervention over the control for effort, domain knowledge, or elaboration-based strategy use.
The subject comparison was also null: intervention effects did not differ significantly between physics and English. Exploratory analyses offered a more useful design clue. Students who interacted more meaningfully with the GPT tended to show more sustained interest and higher domain-specific post-test scores. Teacher observations suggested that some students treated the chatbot primarily as an answer source rather than engaging with its motivational or explanatory scaffolds. Requiring a minimal interaction was therefore not equivalent to securing productive learning dialogue.
Several limitations matter. Thirty-five percent of students had no post-test data, although the authors used intention-to-treat analyses, multiple imputation, and complier-effect checks. The convenience sample came from one German state. The intervention was brief, and some measured strategies may need longer practice before change becomes visible. Standard ChatGPT was also a strong control, so the study cannot establish how any of the GPT conditions compare with equivalent lessons without a chatbot. The paper consequently supports restraint rather than a simple claim that pedagogical prompts work or fail.
For AIEDHK, the practical lesson is that an educational system cannot rely on the presence of a theory-informed prompt. Designers must align the prompt, learner expectations, activity, response demands, and outcome measure. A useful school pilot would teach students why the agent asks them to reflect or explain, examine the substance of their interaction rather than message counts, and include independent assessments of knowledge and strategy use. Longer studies should test retention, transfer, subgroup differences, and gradual fading of support. The paper's strongest contribution is showing that meaningful engagement is an implementation condition: a well-written prompt has little educational value if learners continue to use the system only to obtain answers.


