构建首个评估大模型儿童安全风险的基准,聚焦早期潜在危害。
CAREBench: A Child-Safety Risk Benchmark for Language Models

- 设计500个涵盖12类风险的非敏感测试用例,评估模型识别与应对能力。
- 7个前沿模型失败率2%至58%,不同风险类型表现差异显著。
- 面向开发者提供可落地的安全评估工具,适用于AI伦理与儿童保护场景。
如何评估前沿人工智能系统能否在实际伤害发生前识别儿童安全风险?现有儿童安全评估主要针对儿童性虐待内容,但许多风险始于更早阶段:如模型协助成人操纵、伪装、追踪或孤立未成年人,以及回应中加深儿童对AI的情感依赖而非引导其寻求人类支持。我们提出CAREBench(儿童人工智能风险评估),一个用于评估此类上游儿童安全风险的语言模型基准。CAREBench包含500个跨12类风险的提示,包括诱骗与关系操控、欺骗与伪装、监控与隐私、勒索与性侵犯、AI拟人化、情感依赖及心理疾病敏感性。该基准基于家长和临床专家的响应标注,不包含显性虐待内容或图像;而是评估模型是否能识别、拒绝、缓和或引导高风险互动,防止伤害升级。在七个前沿模型上评估发现,失败率介于2%至58%之间,且不同风险类别表现各异。CAREBench为大语言模型开发者提供了一个负责任的评估框架,帮助识别并修补儿童安全策略中的漏洞。
原文摘要 · Abstract (English)
How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations focus on child sexual abuse material, yet many child-safety failures begin earlier: in model assistance that helps adults manipulate, impersonate, profile, or isolate minors, and in model responses that deepen children's emotional dependence on AI systems rather than redirecting them toward human support. We introduce CAREBench (Child AI Risk Evaluation), a benchmark to assess such upstream child-safety risks in language models. CAREBench contains 500 prompts spanning twelve risk categories, including grooming and relationship engineering, deception and impersonation, surveillance and privacy, sextortion and sexual abuse, AI anthropomorphization, emotional dependency, and mental illness sensitivity. Developed with response annotations from parents and clinicians, the benchmark excludes explicit abuse material and imagery; instead, it evaluates whether models recognize, refuse, de-escalate, or redirect risky interactions before harm becomes overt. Evaluating seven frontier models on our benchmark, we find failure rates ranging from 2% to 58%, with failure patterns that vary across risk categories. CAREBench provides a responsibly scoped evaluation for LLM developers to identify and close gaps in child safety policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。