Back to Research News
A child-safety research team studies branching multi-turn dialogue traces behind a protected classroom observation window without showing harmful content
Conference PaperConference paper20255 Aug 2026· 2 min

Child-specific multi-turn red teaming found safety gaps that adult baselines and single-turn tests missed

Prasanjit Rath, Hari Shrawgi, Parag Agrawal, Sandipan Dandapat

NAACL 2025 Industry Track

500-word summary

A child-safety research team studies branching multi-turn dialogue traces behind a protected classroom observation window without showing harmful content

Listen to the paper summary

Audio summary

0:00/0:00

Rath and colleagues argue that a general safety score is not enough for systems used around children. Their NAACL Industry Track paper develops a child-harm taxonomy and synthetic Child User Models, then uses them to red-team six language-model snapshots. The study does not involve real children. Instead, it constructs 560 synthetic child personas and prompts, creates matched adult baselines and runs conversations for up to five turns. Mistral-7B serves as the automated adversarial model and GPT-4o as the automated judge.

The taxonomy begins with 12 broad child-risk categories and is divided into 14 categories for the experiments. The primary defect measure asks whether a conversation contains at least one response judged harmful. Even the lowest reported family-level defect rate, for the Llama family in this evaluation, was 29.6%. The authors find especially large gaps between child and adult personas for sexual content, at 75.4% versus 16.7%; regulated goods and services, at 71.3% versus 30.0%; illegal activities, at 46.7% versus 9.2%; and education-related harm, at 23.3% versus 8.1%.

Conversation length changes what the benchmark detects. Among first harmful responses, 48.12% appeared in the third turn, compared with 25.25% in the first. A model may refuse an obvious first request but become unsafe as a fictional scenario, personal disclosure or sequence of follow-up questions develops. For schools and developers, that finding supports testing realistic dialogue paths rather than only a list of isolated prohibited prompts.

The numerical results need strict boundaries. Synthetic personas cannot represent the full diversity of children's language, development, disability, culture or circumstances. Both the attacker and judge are models, so their strategies and labels can be biased. The evaluation is English-only, ends after five turns and uses early-2025 model snapshots whose safety tuning may later change. The model ranking is not evidence about current products, and the defect rates are not estimates of how often real children experience harm.

The paper is nevertheless useful as an evaluation design. A child-facing or child-adjacent service can define age-specific risks, generate authorized synthetic scenarios, test several turns, include benign near-neighbours to measure over-refusal and have trained humans review a sample. Tests should include requests that are harmless for adults but developmentally inappropriate for children, as well as moments when the model should direct a learner toward a parent, teacher, counsellor or emergency resource.

For Hong Kong schools, vendor safeguards should be one layer within a wider system. Schools need age-appropriate accounts, clear permitted uses, staff escalation, incident reporting and regular re-testing after model updates. The study's lasting contribution is not a league table; it is the warning that adult safety baselines and one-turn checks can miss risks that emerge through a child's continuing conversation.

Evaluation teams should also document which prompts were excluded, how judges disagreed and how harmful examples are protected from unnecessary exposure. Transparent methods let schools compare releases without circulating unsafe content or overstating a benchmark's precision. Independent child-safety experts should review the protocol before deployment decisions.

Related papers

A diverse group of computing students and an instructor compare AI tutoring responses with a prerequisite map and scaffolded Python exercises
Conference Paper2026
Conference Paper 86

Pedagogical-fit feedback improved 51 of 62 weak AI-tutor responses, but scaffolding gains created trade-offs

Benjamin Barlog, Hudson Craig, Zedong Peng

IEEE 27th International Conference on Information Reuse and Integration for Data Science (IRI 2026)

Barlog, Craig and Peng evaluated ChatGPT, Gemini, Gemma 4 and Qwen 3 across 240 introductory-programming tutoring scenarios with a six-part Pedagogical Suitability Index. Baseline scores differed modestly, while targeted feedback improved 51 of 62 weak cases. The largest gain was scaffolding, but prerequisite ordering and Bloom-level alignment sometimes declined, showing why tutoring quality needs multiple measures rather than one composite score.

AI tutorspedagogical fitscaffolding
Read 500-word summary
A diverse university learning team reviews a browser research trail, a coding workflow and teacher-controlled study materials in a bright campus lab
Industry10 Aug 2026
Industry 87

Product news: OpenAI retires Atlas while Claude Code makes auto mode the default, turning agent handoffs into a learning-design issue

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: OpenAI ended Atlas on August 9 and is moving browser-based agent work into ChatGPT and Codex, while Anthropic will make classifier-governed auto mode the Claude Code default for Pro, Max and Team sessions. Gemini for Education supplies the learning-purpose comparison: teach, learn and work with managed data protections. Together, the products make handoffs, permissions and evidence part of AI literacy.

product newsChatGPT browser agentsClaude Code auto mode
Read 500-word summary
Academic cover for a position paper on ChatGPT and large language models in education
Review2023
Review c3c9681c-ca00-4df2-aa39-490f775df4fb

ChatGPT for good? On opportunities and challenges of large language models for education

Enkelejda Kasneci, Kathrin Sessler, Stefan Kuechemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Guennemann, Eyke Huellermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Juergen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuehn, Gjergji Kasneci

Learning and Individual Differences

A widely cited position paper that balances the educational opportunities of large language models with risks around bias, privacy, assessment, and teacher guidance.

large language modelsChatGPTteacher support
Read 500-word summary