← 返回研究新闻
A child-safety research team studies branching multi-turn dialogue traces behind a protected classroom observation window without showing harmful content
会议论文会议论文20252026年8月5日· 8 min

Child-specific multi-turn red teaming found safety gaps that adult baselines and single-turn tests missed

Prasanjit Rath, Hari Shrawgi, Parag Agrawal, Sandipan Dandapat

NAACL 2025 Industry Track

500 字摘要

A child-safety research team studies branching multi-turn dialogue traces behind a protected classroom observation window without showing harmful content

Rath and colleagues argue that a general safety score is not enough for systems used around children. Their NAACL Industry Track paper develops a child-harm taxonomy and synthetic Child User Models, then uses them to red-team six language-model snapshots. The study does not involve real children. Instead, it constructs 560 synthetic child personas and prompts, creates matched adult baselines and runs conversations for up to five turns. Mistral-7B serves as the automated adversarial model and GPT-4o as the automated judge.

The taxonomy begins with 12 broad child-risk categories and is divided into 14 categories for the experiments. The primary defect measure asks whether a conversation contains at least one response judged harmful. Even the lowest reported family-level defect rate, for the Llama family in this evaluation, was 29.6%. The authors find especially large gaps between child and adult personas for sexual content, at 75.4% versus 16.7%; regulated goods and services, at 71.3% versus 30.0%; illegal activities, at 46.7% versus 9.2%; and education-related harm, at 23.3% versus 8.1%.

Conversation length changes what the benchmark detects. Among first harmful responses, 48.12% appeared in the third turn, compared with 25.25% in the first. A model may refuse an obvious first request but become unsafe as a fictional scenario, personal disclosure or sequence of follow-up questions develops. For schools and developers, that finding supports testing realistic dialogue paths rather than only a list of isolated prohibited prompts.

The numerical results need strict boundaries. Synthetic personas cannot represent the full diversity of children's language, development, disability, culture or circumstances. Both the attacker and judge are models, so their strategies and labels can be biased. The evaluation is English-only, ends after five turns and uses early-2025 model snapshots whose safety tuning may later change. The model ranking is not evidence about current products, and the defect rates are not estimates of how often real children experience harm.

The paper is nevertheless useful as an evaluation design. A child-facing or child-adjacent service can define age-specific risks, generate authorized synthetic scenarios, test several turns, include benign near-neighbours to measure over-refusal and have trained humans review a sample. Tests should include requests that are harmless for adults but developmentally inappropriate for children, as well as moments when the model should direct a learner toward a parent, teacher, counsellor or emergency resource.

For Hong Kong schools, vendor safeguards should be one layer within a wider system. Schools need age-appropriate accounts, clear permitted uses, staff escalation, incident reporting and regular re-testing after model updates. The study's lasting contribution is not a league table; it is the warning that adult safety baselines and one-turn checks can miss risks that emerge through a child's continuing conversation.

Evaluation teams should also document which prompts were excluded, how judges disagreed and how harmful examples are protected from unnecessary exposure. Transparent methods let schools compare releases without circulating unsafe content or overstating a benchmark's precision. Independent child-safety experts should review the protocol before deployment decisions.

相关论文

三位教育与软件同事在明亮的大学设计工作室审查图解教材卡、注释图表和数字原型
政策 / 伦理2026年9月7日
政策 / 伦理 113

评论:Fable 5.1 把更长程的 AI 工作带进 AIED,教育验证更显重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 于 2026 年 9 月 1 日发布 Claude Fable 5.1,强化长程编程与知识工作能力,并降低缓存读取价格。AIED 的机会,是加快从教学构想到可审查原型与研究分析的循环;真正的考验,是能否把速度转化为更好的教学与可信证据,同时计入总成本、数据条件和人工审核。

产品新闻评论Claude Fable 5.1
阅读 500 字摘要 →
A racially diverse secondary-school class uses teacher-selected course materials to build a study guide while an educator reviews the activity on a classroom display
政策 / 伦理2026年8月18日
政策 / 伦理 105

Product news: Gemini in Classroom expands to students of all ages with course-grounded study prompts

Google Workspace

AI Product and Learning Report

Product news: Google is expanding Gemini in Classroom to eligible K-12 and higher-education students of all ages, with contextual prompts that can use a selected class, assignment instructions, and curriculum materials. Web rollout began August 10 and mobile rollout August 17. Admin controls remain decisive, and Google advises reviewing generated output, so access should be paired with age-appropriate teaching, privacy checks, and learning evidence.

product newsGemini in ClassroomK-12 AI
阅读 500 字摘要 →
Academic cover for a position paper on ChatGPT and large language models in education
综述2023
综述 c3c9681c-ca00-4df2-aa39-490f775df4fb

ChatGPT for good? On opportunities and challenges of large language models for education

Enkelejda Kasneci, Kathrin Sessler, Stefan Kuechemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Guennemann, Eyke Huellermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Juergen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuehn, Gjergji Kasneci

Learning and Individual Differences

A widely cited position paper that balances the educational opportunities of large language models with risks around bias, privacy, assessment, and teacher guidance.

large language modelsChatGPTteacher support
阅读 500 字摘要 →