为7-11岁儿童设计了大模型安全评估基准,发现显式年龄提示可显著提升响应安全性。
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models

- 基于儿童发展心理学构建评分体系,评估大模型对儿童提问的安全响应。
- 隐式提示使安全得分提升9%-47%,显式年龄指令再增10%-30%。
- 适用于儿童友好型AI开发与多语言文化场景的安全优化参考。
儿童日益接触大型语言模型(LLMs),可能面临不适宜其发展阶段的回应,亟需具备年龄敏感性、安全性和边界感的防护机制。现有大模型安全评估主要聚焦有害内容规避,缺乏对儿童面向安全的针对性。本文提出KIDBench,一个针对7-11岁儿童的基准测试,采用发展心理学指导的“大模型作为评判者”评分标准,包含十类真实儿童提问,涵盖单轮与多轮对话模拟。通过对比无提示、隐式线索(暗示儿童身份)和显式年龄指令三种方式发现:隐式线索使安全得分提升9%-47%,显式年龄指令进一步带来10%-30%的增益。跨语言与文化评估显示不同语言和地区间安全表现差异明显。多轮模拟表明,儿童相关回应质量在后续对话中可下降6%-24%。除评估外,本文还推出KIDGuardLlama(安全评测模型)与KIDLlama(儿童导向响应模型),展示如何利用KIDBench推动更安全的儿童交互人工智能。
原文摘要 · Abstract (English)
Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require age-sensitive safety, guidance, and boundaries. Existing LLM safety evaluations largely focus on harmful-content avoidance and do not explicitly target child-facing safety. We introduce KIDBench, a benchmark for evaluating child-facing LLM safety for ages 7-11 using a developmental-psychology-grounded LLM-as-a-Judge rubric. KIDBench contains realistic child queries across ten categories, with single-turn prompts and multi-turn child-actor simulations. We compare no-cues prompts with no child context, implicit-cues prompts that suggest a child speaker, and explicit age instructions. Implicit-cues improve scores by 9-47% across models, while explicit age adds a further 10-30% gain. Cross-lingual and cultural evaluations show uneven safety behavior across languages and country contexts. Multi-turn simulations show that child-facing response quality can degrade by 6-24% from the first to worst turn. Beyond evaluation, we introduce KIDGuardLlama, a child-safety evaluator, and KIDLlama, a child-oriented response model, showing how KIDBench supports safer child-facing AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。