
Students' essays supported a scaffold-don't-substitute principle, but not a causal harm estimate
Lucile Favero, Juan Antonio Pérez-Ortiz, Tanja Käser, Nuria Oliver
ACM AI Leadership Summit
500-word summary

Listen to the paper summary
Audio summary
Lucile Favero, Juan Antonio Pérez-Ortiz, Tanja Käser and Nuria Oliver argue that educational AI becomes harmful when it substitutes for the human capacities education is intended to develop. Their framework connects four dimensions: cognition, agency, emotional wellbeing, and ethics. Cognitive offloading can reduce effort; reduced effort can weaken a learner's sense of control; dependence and uncertainty can affect wellbeing; and these changes can intensify questions about responsibility and fairness. The authors describe this as a self-reinforcing cycle rather than four isolated risks.
The paper grounds the framework in an exploratory analysis of 49 International Baccalaureate argumentative essays about AI's impact. Eighty percent of the essays reported that reliance on AI reduces thinking. Students also described a preferred alternative: systems that withhold immediate answers, prompt recall, ask questions, and encourage reflection. The authors connect these requests with established learning-science principles and summarize the resulting design stance as scaffold, do not substitute.
Student voice is a strength of the study. Learners are often treated as recipients of AI policy rather than contributors to product design. Their essays reveal concerns and desired interactions in their own educational context. Still, the number 80 percent must be interpreted narrowly. The sample is small, comes from a particular programme, and consists of essays written for an argument task. It does not measure actual tool use, cognitive change, achievement, mental health, or the prevalence of an effect across students. Coding self-reports cannot establish that AI caused the harms described.
The proposed design principle is therefore best treated as a hypothesis and a requirement for testing. A scaffolded tutor might ask a learner to retrieve an idea before showing a hint, request a prediction before a simulation, or provide a partial step followed by an explanation prompt. A substitutive tool might produce the completed essay or solution immediately. Researchers can compare these designs on unaided retention, transfer, help-seeking, confidence calibration, frustration, agency, accessibility, and time. They should also examine whether withholding help disadvantages learners who need accommodations or foundational instruction.
Teachers can apply the principle without banning AI. They can specify phases in which learners first attempt, explain, or retrieve; allow graduated hints; require source checks; and assess a later task without assistance. Students should help define when support feels productive and when it feels controlling or evasive. Product teams should expose hint policies and let educators align them with age, subject, and learner needs.
For AIEDHK, the paper supplies a memorable orientation while modelling evidential restraint. The student essays justify taking substitution risks seriously, but they do not quantify causal harm. The next step is participatory, prospective evaluation of designs that preserve effort without denying appropriate support. Educational AI should be judged by the capacities learners retain and can exercise independently, not by the amount of work the system completes on their behalf. Designers should report when scaffolds frustrate, exclude, or delay learners as carefully as they report successful reflection, because support quality depends on responsive adjustment.


