← 返回研究新聞
A diverse teacher team reviews a staged AI-assisted computer-networking lesson plan with learning objectives, instructional units and classroom activities
期刊論文同行評審研究20262026年7月16日· 8 min

Pedagogy-grounded prompts improved DeepSeek-generated lesson plans, but classroom learning was not tested

Yinan Lu, Weinuo Li, Yue Cai

Systems

500 字摘要

A diverse teacher team reviews a staged AI-assisted computer-networking lesson plan with learning objectives, instructional units and classroom activities

Lu, Li and Cai test whether pedagogical structure can improve large-language-model lesson plans beyond a single general prompt. Their framework divides planning into three dependent stages: measurable learning-objective design, instructional-unit design and teaching-activity design. Bloom's taxonomy guides the objectives. Either ACT-R phases or Gagné's instructional-event structure guides the units, while problem-chain theory can guide activities. Outputs from one stage become inputs to the next, and the design allows a teacher to inspect or revise an intermediate result before generation continues.

The experiment used three DeepSeek models: R1, V3 and R1-Distill-Qwen-32B. Researchers crossed them with five prompting strategies across ten authentic topics drawn from a Computer Networks course, generating 15 plans per topic and 150 plans in total. Generation used a fixed temperature of 0.7 and a fixed random seed. The comparison included a naive baseline plus four theory-grounded variants that combined either ACT-R or Gagné unit structures with basic or problem-chain activity design.

Evaluation covered both structural completeness and functional quality. An LLM judge scored the full 150-plan corpus for coverage of Gagné's nine events and a 12-indicator functional rubric. Five university computer-science teachers, blind to model and strategy, rated a stratified sample of 30 plans after calibration. Across the complete automated evaluation, every theory-based strategy exceeded the naive baseline for every model and raised event coverage above 90%. The smaller 32B model showed the largest structural rise, from 71.6% under baseline to 98.7% in its strongest condition.

Human ratings supported the direction but narrowed the certainty. Average event coverage in the human sample was 74.1% for baseline plans and 89.8% to 93.5% for theory-grounded strategies. The paper reports functional-quality improvements of up to 17.3% under the LLM judge and 53.6% under human rating. Gagné-based structure performed better than ACT-R under base conditions, while problem-chain guidance particularly benefited the ACT-R route. The LLM judge nevertheless scored plans systematically higher than humans, so automated absolute scores should not be treated as objective teaching quality.

The study evaluates generated documents, not live co-design or classroom outcomes. All topics came from one university Computer Networks course, all models belonged to one model family, and only 30 of 150 plans received human validation. Teachers were expert evaluators rather than participants whose real planning time, revisions or acceptance decisions were studied. No students received the lessons, and the design measured neither engagement, attainment nor transfer. It is therefore accurate to report improved plan coverage and rubric scores, but not improved teaching or learning.

For Hong Kong, the three-stage workflow is a testable professional-learning scaffold. Teachers could compare a one-shot draft with a staged draft aligned to local curricula, revise each intermediate output and record review time, factual errors, accessibility and language quality in Cantonese, Chinese and English. A classroom pilot should then score independent student work and teacher workload rather than assuming a complete-looking plan is effective. The paper's practical contribution is a structured prompting hypothesis; local educators still supply curriculum judgment, safety review and the evidence needed to decide whether the resulting lesson works.

相關論文

大學生向教師解釋幾何作圖,同學在旁思考,桌上的電腦展示相關數碼圖解
政策 / 倫理2026年9月7日
政策 / 倫理 112

評論:Astra 的 AGI 主張,讓教育更需要看見人的真實學習

AIED.HK Editorial

AI Product News Commentary

OpenAI 於 2026 年 9 月 3 日發布 GPT-6 Astra,並引發關於 AGI 是否已經到來的討論。本文將這一說法視為需要歸屬的主張,而非已確立的共識。教育眼前的挑戰,是分清 AI 能產出甚麼,以及學習者能獨立解釋、質疑和遷移甚麼,再以更強的代理能力支援真正的學習。

產品新聞評論GPT-6 Astra
閱讀 500 字摘要 →
三位教育與軟件同事在明亮的大學設計工作室審查圖解教材卡、註釋圖表和數碼原型
政策 / 倫理2026年9月7日
政策 / 倫理 113

評論:Fable 5.1 把更長程的 AI 工作帶進 AIED,教育驗證更顯重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 於 2026 年 9 月 1 日發布 Claude Fable 5.1,強化長程編程與知識工作能力,並降低快取讀取價格。AIED 的機會,是加快從教學構想到可審查原型與研究分析的循環;真正的考驗,是能否把速度轉化為更好的教學與可信證據,同時計入總成本、資料條件和人工審核。

產品新聞評論Claude Fable 5.1
閱讀 500 字摘要 →
Four diverse university students practise prompting and source checking with an instructor at a library learning table
期刊論文2026
期刊論文 52

A 90-minute GenAI literacy course improved knowledge, prompting, source checking and self-efficacy across 65 university sections

Allison E. Connell Pensky, Lydia E. Eckstein, Michael C. Melville, Laura O. Pottmeyer, Zach Mineroff, Avi Chawla, Judy Brooks, Chad Hershock, Marsha C. Lovett

Computers & Education

In a large experiment involving 1,368 undergraduate and graduate students across 65 university course sections, a 90-minute asynchronous GenAI learning module improved knowledge of how the technology works, prompt-engineering performance, fact- and source-checking, and self-efficacy. It did not improve critical evaluation of bias, showing that short foundational training needs deeper practice for responsible judgment.

generative AI literacyrandomized experimenthigher education
閱讀 500 字摘要 →