← 返回研究新聞
Three diverse university science students compare three diagram drafts beside a text-free laptop feedback view linked by one cyan revision path
期刊論文同行評審研究20262026年7月27日· 10 min

Human-centered GenAI feedback design in higher education: a multisite experiment on direct, reflective, and hybrid approaches to scientific argumentation

Huseyin Ates

International Journal of Educational Technology in Higher Education

500 字摘要

Three diverse university science students compare three diagram drafts beside a text-free laptop feedback view linked by one cyan revision path

Ates tests a question that is more useful than whether generative AI can produce feedback: which feedback design helps students revise now and perform later without AI? The 2026 open-access study used a multisite, cluster-randomized, longitudinal field experiment in introductory biology, chemistry, and physics courses. It compared peer feedback, direct GenAI feedback, reflective GenAI feedback, and a hybrid sequence of self-evaluation, peer feedback, and GenAI critique.

The analytic sample included 1,176 first-year undergraduates in 48 course sections across four universities. Sections, rather than individual students, were assigned to conditions. The four groups were broadly comparable at baseline on demographics, prior knowledge, argumentation, achievement, feedback literacy, and previous GenAI experience. Attrition did not differ significantly by condition, and implementation audits indicated that 95.8 percent of applicable instructional steps were delivered as planned.

Students completed three cycles of drafting and revision around scientific arguments. Quality was scored on claims, relevance and sufficiency of evidence, coherence of reasoning, and treatment of limitations or alternative explanations. The study also measured conceptual learning, feedback uptake, self-regulated learning during revision, and a delayed transfer task completed individually under supervision without the GenAI tool or internet-enabled devices.

Direct GenAI feedback outperformed peer feedback on immediate argument-quality gain. Yet both reflective and hybrid feedback produced stronger immediate gains than direct AI feedback, and the hybrid condition had the highest adjusted mean. The hybrid-versus-reflective difference was not statistically significant. Revision-depth analyses followed the same pattern, suggesting that improvements were not limited to surface editing.

The differences became more educationally consequential beyond the revised product. Hybrid feedback significantly outperformed direct GenAI feedback on conceptual learning. The reflective contrast was positive but did not remain statistically significant after adjustment. On delayed AI-free transfer, both reflective and hybrid conditions significantly outperformed direct GenAI feedback, with no significant difference between them.

Process evidence helps explain the pattern. Reflective and hybrid designs produced higher feedback uptake and self-regulated learning than direct GenAI feedback. Multilevel mediation models found significant indirect pathways through uptake, self-regulation, and their sequence for argument gain and delayed transfer. The results support the interpretation that students learned more when the design required them to judge and work with feedback rather than simply receive a polished critique.

The study is unusually strong for educational GenAI research because it spans institutions, randomizes clusters, checks implementation, uses manually scored disciplinary work, and includes a delayed AI-free outcome. It also preserves important boundaries: submitted work had to remain the student's own, students received guidance on ethical use and data-entry limits, and monitoring was restricted to the study platform.

Limits remain. Cluster assignment leaves only 48 randomized units, the author conducted all major study functions, and the intervention focused on first-year science argumentation. The paper cannot establish that the same sequence will transfer to other disciplines, age groups, commercial tools, or longer periods. Self-regulated learning was partly measured through self-report, even though revision traces provided complementary evidence.

For Hong Kong higher education, the strongest design implication is sequencing. Ask students to assess their draft against criteria before seeing AI critique, incorporate peer evidence, require a rationale for accepted and rejected suggestions, and test later on a new task without AI. The study suggests that feedback becomes learning when students retain evaluative judgment and ownership, not when the system merely produces more comments.

相關論文

大學生向教師解釋幾何作圖,同學在旁思考,桌上的電腦展示相關數碼圖解
政策 / 倫理2026年9月7日
政策 / 倫理 112

評論:Astra 的 AGI 主張,讓教育更需要看見人的真實學習

AIED.HK Editorial

AI Product News Commentary

OpenAI 於 2026 年 9 月 3 日發布 GPT-6 Astra,並引發關於 AGI 是否已經到來的討論。本文將這一說法視為需要歸屬的主張,而非已確立的共識。教育眼前的挑戰,是分清 AI 能產出甚麼,以及學習者能獨立解釋、質疑和遷移甚麼,再以更強的代理能力支援真正的學習。

產品新聞評論GPT-6 Astra
閱讀 500 字摘要 →
三位教育與軟件同事在明亮的大學設計工作室審查圖解教材卡、註釋圖表和數碼原型
政策 / 倫理2026年9月7日
政策 / 倫理 113

評論:Fable 5.1 把更長程的 AI 工作帶進 AIED,教育驗證更顯重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 於 2026 年 9 月 1 日發布 Claude Fable 5.1,強化長程編程與知識工作能力,並降低快取讀取價格。AIED 的機會,是加快從教學構想到可審查原型與研究分析的循環;真正的考驗,是能否把速度轉化為更好的教學與可信證據,同時計入總成本、資料條件和人工審核。

產品新聞評論Claude Fable 5.1
閱讀 500 字摘要 →
A lecturer and two university students inspect ranked learning tools, separate cloud and local plugin cards, and a review ledger in a bright computing studio
政策 / 倫理2026年8月23日
政策 / 倫理 111

Product news: ChatGPT plugin ranking and Claude Code 2.1.239 make tool selection and workspace boundaries inspectable

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: ChatGPT now ranks plugin recommendations partly by continued use after installation and adds more time-aware answers, while Claude Code 2.1.239 distinguishes cloud-synced plugins from local installations and makes a data-residency cost premium visible. Gemini for Education supplies the institutional purpose boundary across teaching, learning and work. Together, the updates make tool selection, context, cost and human review part of AI workflow literacy.

product newsChatGPT pluginsClaude Code 2.1.239
閱讀 500 字摘要 →