
Child-specific multi-turn red teaming found safety gaps that adult baselines and single-turn tests missed
Prasanjit Rath, Hari Shrawgi, Parag Agrawal, Sandipan Dandapat
NAACL 2025 Industry Track
Rath and colleagues created 560 synthetic child personas and matched adult baselines to red-team six language-model snapshots over five-turn conversations. The benchmark found substantially higher defect rates for child scenarios in several categories and showed that many failures emerged only after the dialogue developed.
