Back to Research News
A child-safety research team studies branching multi-turn dialogue traces behind a protected classroom observation window without showing harmful content
Conference PaperConference paper20255 Aug 2026· 2 min

Child-specific multi-turn red teaming found safety gaps that adult baselines and single-turn tests missed

Prasanjit Rath, Hari Shrawgi, Parag Agrawal, Sandipan Dandapat

NAACL 2025 Industry Track

500-word summary

A child-safety research team studies branching multi-turn dialogue traces behind a protected classroom observation window without showing harmful content

Listen to the paper summary

Audio summary

0:00/0:00

Rath and colleagues argue that a general safety score is not enough for systems used around children. Their NAACL Industry Track paper develops a child-harm taxonomy and synthetic Child User Models, then uses them to red-team six language-model snapshots. The study does not involve real children. Instead, it constructs 560 synthetic child personas and prompts, creates matched adult baselines and runs conversations for up to five turns. Mistral-7B serves as the automated adversarial model and GPT-4o as the automated judge.

The taxonomy begins with 12 broad child-risk categories and is divided into 14 categories for the experiments. The primary defect measure asks whether a conversation contains at least one response judged harmful. Even the lowest reported family-level defect rate, for the Llama family in this evaluation, was 29.6%. The authors find especially large gaps between child and adult personas for sexual content, at 75.4% versus 16.7%; regulated goods and services, at 71.3% versus 30.0%; illegal activities, at 46.7% versus 9.2%; and education-related harm, at 23.3% versus 8.1%.

Conversation length changes what the benchmark detects. Among first harmful responses, 48.12% appeared in the third turn, compared with 25.25% in the first. A model may refuse an obvious first request but become unsafe as a fictional scenario, personal disclosure or sequence of follow-up questions develops. For schools and developers, that finding supports testing realistic dialogue paths rather than only a list of isolated prohibited prompts.

The numerical results need strict boundaries. Synthetic personas cannot represent the full diversity of children's language, development, disability, culture or circumstances. Both the attacker and judge are models, so their strategies and labels can be biased. The evaluation is English-only, ends after five turns and uses early-2025 model snapshots whose safety tuning may later change. The model ranking is not evidence about current products, and the defect rates are not estimates of how often real children experience harm.

The paper is nevertheless useful as an evaluation design. A child-facing or child-adjacent service can define age-specific risks, generate authorized synthetic scenarios, test several turns, include benign near-neighbours to measure over-refusal and have trained humans review a sample. Tests should include requests that are harmless for adults but developmentally inappropriate for children, as well as moments when the model should direct a learner toward a parent, teacher, counsellor or emergency resource.

For Hong Kong schools, vendor safeguards should be one layer within a wider system. Schools need age-appropriate accounts, clear permitted uses, staff escalation, incident reporting and regular re-testing after model updates. The study's lasting contribution is not a league table; it is the warning that adult safety baselines and one-turn checks can miss risks that emerge through a child's continuing conversation.

Evaluation teams should also document which prompts were excluded, how judges disagreed and how harmful examples are protected from unnecessary exposure. Transparent methods let schools compare releases without circulating unsafe content or overstating a benchmark's precision. Independent child-safety experts should review the protocol before deployment decisions.

Related papers

Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory
Journal Paper2026
Journal Paper 54

Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward

Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi

arXiv working paper

In a preregistered randomized online experiment with 1,174 Argentine adults, GPT-4.1 assistance raised workplace-style problem-solving performance for both education groups and reduced the baseline gap from 0.548 to 0.139 standard deviations. Lower-education participants retained a modest gain after AI was removed, but stronger follow-up performance appeared when intensive assistance was paired with sustained human effort.

generative AIrandomized experimenteducation inequality
Read 500-word summary
A Black female lecturer and two diverse university students review an audio transcript, curriculum binder and organized learning cards in a media studio
Industry9 Aug 2026
Industry 55

Product news: GPT Transcribe, Claude memory and Gemini Classroom make learning context persistent

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: OpenAI released GPT Transcribe and GPT Live Transcribe for file and streaming speech, Anthropic changed Claude memory into categorized entries that update across conversations, and Google is connecting Gemini learning activities to teacher-selected Classroom materials. Together, the products make consent, correction and purposeful forgetting central to educational AI design.

product newsGPT TranscribeClaude memory
Read 500-word summary
Academic cover for a position paper on ChatGPT and large language models in education
Review2023
Review c3c9681c-ca00-4df2-aa39-490f775df4fb

ChatGPT for good? On opportunities and challenges of large language models for education

Enkelejda Kasneci, Kathrin Sessler, Stefan Kuechemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Guennemann, Eyke Huellermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Juergen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuehn, Gjergji Kasneci

Learning and Individual Differences

A widely cited position paper that balances the educational opportunities of large language models with risks around bias, privacy, assessment, and teacher guidance.

large language modelsChatGPTteacher support
Read 500-word summary