← 返回研究新闻
A diverse teacher team reviews a staged AI-assisted computer-networking lesson plan with learning objectives, instructional units and classroom activities
期刊论文同行评审研究20262026年7月16日· 8 min

Pedagogy-grounded prompts improved DeepSeek-generated lesson plans, but classroom learning was not tested

Yinan Lu, Weinuo Li, Yue Cai

Systems

500 字摘要

A diverse teacher team reviews a staged AI-assisted computer-networking lesson plan with learning objectives, instructional units and classroom activities

Lu, Li and Cai test whether pedagogical structure can improve large-language-model lesson plans beyond a single general prompt. Their framework divides planning into three dependent stages: measurable learning-objective design, instructional-unit design and teaching-activity design. Bloom's taxonomy guides the objectives. Either ACT-R phases or Gagné's instructional-event structure guides the units, while problem-chain theory can guide activities. Outputs from one stage become inputs to the next, and the design allows a teacher to inspect or revise an intermediate result before generation continues.

The experiment used three DeepSeek models: R1, V3 and R1-Distill-Qwen-32B. Researchers crossed them with five prompting strategies across ten authentic topics drawn from a Computer Networks course, generating 15 plans per topic and 150 plans in total. Generation used a fixed temperature of 0.7 and a fixed random seed. The comparison included a naive baseline plus four theory-grounded variants that combined either ACT-R or Gagné unit structures with basic or problem-chain activity design.

Evaluation covered both structural completeness and functional quality. An LLM judge scored the full 150-plan corpus for coverage of Gagné's nine events and a 12-indicator functional rubric. Five university computer-science teachers, blind to model and strategy, rated a stratified sample of 30 plans after calibration. Across the complete automated evaluation, every theory-based strategy exceeded the naive baseline for every model and raised event coverage above 90%. The smaller 32B model showed the largest structural rise, from 71.6% under baseline to 98.7% in its strongest condition.

Human ratings supported the direction but narrowed the certainty. Average event coverage in the human sample was 74.1% for baseline plans and 89.8% to 93.5% for theory-grounded strategies. The paper reports functional-quality improvements of up to 17.3% under the LLM judge and 53.6% under human rating. Gagné-based structure performed better than ACT-R under base conditions, while problem-chain guidance particularly benefited the ACT-R route. The LLM judge nevertheless scored plans systematically higher than humans, so automated absolute scores should not be treated as objective teaching quality.

The study evaluates generated documents, not live co-design or classroom outcomes. All topics came from one university Computer Networks course, all models belonged to one model family, and only 30 of 150 plans received human validation. Teachers were expert evaluators rather than participants whose real planning time, revisions or acceptance decisions were studied. No students received the lessons, and the design measured neither engagement, attainment nor transfer. It is therefore accurate to report improved plan coverage and rubric scores, but not improved teaching or learning.

For Hong Kong, the three-stage workflow is a testable professional-learning scaffold. Teachers could compare a one-shot draft with a staged draft aligned to local curricula, revise each intermediate output and record review time, factual errors, accessibility and language quality in Cantonese, Chinese and English. A classroom pilot should then score independent student work and teacher workload rather than assuming a complete-looking plan is effective. The paper's practical contribution is a structured prompting hypothesis; local educators still supply curriculum judgment, safety review and the evidence needed to decide whether the resulting lesson works.

相关论文

大学生向教师解释几何作图,同学在旁思考,桌上的电脑展示相关数字图解
政策 / 伦理2026年9月7日
政策 / 伦理 112

评论:Astra 的 AGI 主张,让教育更需要看见人的真实学习

AIED.HK Editorial

AI Product News Commentary

OpenAI 于 2026 年 9 月 3 日发布 GPT-6 Astra,并引发关于 AGI 是否已经到来的讨论。本文将这一说法视为需要归属的主张,而非已确立的共识。教育眼前的挑战,是分清 AI 能产出什么,以及学习者能独立解释、质疑和迁移什么,再以更强的代理能力支持真正的学习。

产品新闻评论GPT-6 Astra
阅读 500 字摘要 →
三位教育与软件同事在明亮的大学设计工作室审查图解教材卡、注释图表和数字原型
政策 / 伦理2026年9月7日
政策 / 伦理 113

评论:Fable 5.1 把更长程的 AI 工作带进 AIED,教育验证更显重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 于 2026 年 9 月 1 日发布 Claude Fable 5.1,强化长程编程与知识工作能力,并降低缓存读取价格。AIED 的机会,是加快从教学构想到可审查原型与研究分析的循环;真正的考验,是能否把速度转化为更好的教学与可信证据,同时计入总成本、数据条件和人工审核。

产品新闻评论Claude Fable 5.1
阅读 500 字摘要 →
Four diverse university students practise prompting and source checking with an instructor at a library learning table
期刊论文2026
期刊论文 52

A 90-minute GenAI literacy course improved knowledge, prompting, source checking and self-efficacy across 65 university sections

Allison E. Connell Pensky, Lydia E. Eckstein, Michael C. Melville, Laura O. Pottmeyer, Zach Mineroff, Avi Chawla, Judy Brooks, Chad Hershock, Marsha C. Lovett

Computers & Education

In a large experiment involving 1,368 undergraduate and graduate students across 65 university course sections, a 90-minute asynchronous GenAI learning module improved knowledge of how the technology works, prompt-engineering performance, fact- and source-checking, and self-efficacy. It did not improve critical evaluation of bias, showing that short foundational training needs deeper practice for responsible judgment.

generative AI literacyrandomized experimenthigher education
阅读 500 字摘要 →