
ورقة مؤتمر2025
ورقة مؤتمر 66Child-specific multi-turn red teaming found safety gaps that adult baselines and single-turn tests missed
Prasanjit Rath, Hari Shrawgi, Parag Agrawal, Sandipan Dandapat
NAACL 2025 Industry Track
Rath and colleagues created 560 synthetic child personas and matched adult baselines to red-team six language-model snapshots over five-turn conversations. The benchmark found substantially higher defect rates for child scenarios in several categories and showed that many failures emerged only after the dialogue developed.
child safetylarge language modelsred teaming
اقرأ ملخص 500 كلمة →