연구 뉴스로 돌아가기
A diverse group of educators reviews an instructional video storyboard, narration timeline, and multimedia-learning checks at a university media lab
학회 논문Conference paper20262026년 8월 23일· 8 min

Dual gatekeeping improved ratings of AI-generated instructional video without testing student learning

Yearim Kim, Njun Baek, Nojun Kwak

CHI 2026 Workshop on Understanding and Engaging Critical Resistance to AI in Education

500단어 요약

A diverse group of educators reviews an instructional video storyboard, narration timeline, and multimedia-learning checks at a university media lab

Yearim Kim, Njun Baek and Nojun Kwak study a common risk in AI-generated educational video: a polished output can still sequence ideas poorly, include distracting material or misalign narration and visuals. Their PedaCo system treats resistance to an AI draft as a designed part of authoring rather than a failure. It combines educator review before rendering with automated checks after video synthesis.

The first gate operates at the script stage. Educators select criteria from Mayer's Cognitive Theory of Multimedia Learning, including coherence, signaling, redundancy, contiguity, segmenting and pre-training. An AI reviewer flags possible violations by principle, and the educator can reject, revise or regenerate the script. The system presents its comments as possible problems rather than final judgments. Reviewing text upstream is cheaper than repairing a pedagogical error after narration and visuals have been rendered.

The second gate evaluates the finished video on five dimensions: coherence, redundancy, temporal contiguity, modality and image quality. The educator sees the scores and decides whether to accept the video or return to the script. The design intentionally reserves context-sensitive decisions, such as tone and learner appropriateness, for people while using automation for structural features that are easier to measure.

The authors evaluated the workflow in two ways. In a within-subject study, 23 educators briefed on the multimedia-learning principles compared reviewed videos with videos generated without those guidelines across three topics representing causal, abstract and procedural demands. Mean ratings increased from 3.07 to 3.86 on a five-point scale. The paper reports statistically significant improvements across the principles, including large gains in prerequisite sequencing, removal of irrelevant material and overall instructional validity. Participants also rated production efficiency 4.26 out of five.

A separate automated comparison used 14 videos covering seven science and philosophy topics, with two conditions per topic. Coherence rose from 0.646 to 0.729 and temporal contiguity from 0.273 to 0.294; both differences were statistically significant. Modality, redundancy and image quality did not differ significantly. The authors keep those measures as possible regression checks rather than claiming universal improvement.

The evidence is promising but early. The paper is a four-page workshop report, the educator sample is small, and participants were already briefed on the framework used for judging. The study evaluates ratings and proxy metrics, not comprehension, retention, transfer or accessibility for learners. It also does not establish whether repeated review remains manageable during ordinary teaching or whether automated flags work equally well across languages, ages and subjects.

For AIEDHK, the useful contribution is the placement of two explicit acceptance gates. A school or university could require a teacher-reviewed script before generation, then inspect timing, coherence, captions, accessibility and source accuracy on the final media. A stronger trial should compare student learning and teacher workload across workflows, pre-register outcomes, include bilingual materials and report disagreements between reviewers and metrics. PedaCo supports a practical principle: generated teaching media should remain "not yet" until both professional judgment and inspectable checks support its use.

관련 논문

A lecturer and two university students inspect ranked learning tools, separate cloud and local plugin cards, and a review ledger in a bright computing studio
정책 / 윤리2026년 8월 23일
정책 / 윤리 111

Product news: ChatGPT plugin ranking and Claude Code 2.1.239 make tool selection and workspace boundaries inspectable

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: ChatGPT now ranks plugin recommendations partly by continued use after installation and adds more time-aware answers, while Claude Code 2.1.239 distinguishes cloud-synced plugins from local installations and makes a data-residency cost premium visible. Gemini for Education supplies the institutional purpose boundary across teaching, learning and work. Together, the updates make tool selection, context, cost and human review part of AI workflow literacy.

product newsChatGPT pluginsClaude Code 2.1.239
500단어 요약 읽기
A diverse university project team reviews a compact result panel, a gateway diagram, and a verification checklist in a bright computing studio
정책 / 윤리2026년 8월 21일
정책 / 윤리 109

Product news: Claude Code 2.1.237 adds a concise output style and repairs gateway prompt caching

Anthropic

AI Product and Learning Report

Product news: Claude Code 2.1.237 introduces a built-in Concise output style and fixes prompt caching for sessions that use an LLM gateway or custom base URL. The release can reduce narration and repeated processing, but brevity and cache efficiency do not establish correctness. Educational teams should preserve task requirements, evidence, tests, and review notes outside the presentation style.

product newsClaude Code 2.1.237output styles
500단어 요약 읽기
A diverse group of university students and a lecturer examine clustered dialogue cards, a ten-trait matrix and an exam-progress chart in a bright learning analytics studio
도구 / 데이터셋2026
도구 / 데이터셋 94

Principal Trait Analysis linked AI-tutor dialogue patterns to outcomes, but not yet to transferable skills

Hunter McNichols, Kai Du, Andrew Lan

arXiv preprint

McNichols, Du and Lan introduce Principal Trait Analysis, an LLM-assisted pipeline that turns human-AI conversation traces into interpretable behavioral traits. On 1,540 university AI-tutor sessions and 2,774 professional coding-agent sessions, selected traits added explanatory or predictive signal beyond prior performance. Conceptual questioning aligned positively with some exam outcomes, but cross-semester inconsistency, contradictory coefficients and mostly flat temporal patterns mean the traits cannot yet be treated as transferable AI collaboration skills.

Principal Trait AnalysisAI tutoring dialoguehuman-AI collaboration
500단어 요약 읽기