
Dual gatekeeping improved ratings of AI-generated instructional video without testing student learning
Yearim Kim, Njun Baek, Nojun Kwak
CHI 2026 Workshop on Understanding and Engaging Critical Resistance to AI in Education
500-शब्द सार

Yearim Kim, Njun Baek and Nojun Kwak study a common risk in AI-generated educational video: a polished output can still sequence ideas poorly, include distracting material or misalign narration and visuals. Their PedaCo system treats resistance to an AI draft as a designed part of authoring rather than a failure. It combines educator review before rendering with automated checks after video synthesis.
The first gate operates at the script stage. Educators select criteria from Mayer's Cognitive Theory of Multimedia Learning, including coherence, signaling, redundancy, contiguity, segmenting and pre-training. An AI reviewer flags possible violations by principle, and the educator can reject, revise or regenerate the script. The system presents its comments as possible problems rather than final judgments. Reviewing text upstream is cheaper than repairing a pedagogical error after narration and visuals have been rendered.
The second gate evaluates the finished video on five dimensions: coherence, redundancy, temporal contiguity, modality and image quality. The educator sees the scores and decides whether to accept the video or return to the script. The design intentionally reserves context-sensitive decisions, such as tone and learner appropriateness, for people while using automation for structural features that are easier to measure.
The authors evaluated the workflow in two ways. In a within-subject study, 23 educators briefed on the multimedia-learning principles compared reviewed videos with videos generated without those guidelines across three topics representing causal, abstract and procedural demands. Mean ratings increased from 3.07 to 3.86 on a five-point scale. The paper reports statistically significant improvements across the principles, including large gains in prerequisite sequencing, removal of irrelevant material and overall instructional validity. Participants also rated production efficiency 4.26 out of five.
A separate automated comparison used 14 videos covering seven science and philosophy topics, with two conditions per topic. Coherence rose from 0.646 to 0.729 and temporal contiguity from 0.273 to 0.294; both differences were statistically significant. Modality, redundancy and image quality did not differ significantly. The authors keep those measures as possible regression checks rather than claiming universal improvement.
The evidence is promising but early. The paper is a four-page workshop report, the educator sample is small, and participants were already briefed on the framework used for judging. The study evaluates ratings and proxy metrics, not comprehension, retention, transfer or accessibility for learners. It also does not establish whether repeated review remains manageable during ordinary teaching or whether automated flags work equally well across languages, ages and subjects.
For AIEDHK, the useful contribution is the placement of two explicit acceptance gates. A school or university could require a teacher-reviewed script before generation, then inspect timing, coherence, captions, accessibility and source accuracy on the final media. A stronger trial should compare student learning and teacher workload across workflows, pre-register outcomes, include bilingual materials and report disagreements between reviewers and metrics. PedaCo supports a practical principle: generated teaching media should remain "not yet" until both professional judgment and inspectable checks support its use.


