为青少年大模型安全设计新评测基准,发现多个关键风险
SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
- 构建1283个针对儿童发展的对抗性提示,覆盖0-18岁各阶段风险
- 47个模型测试显示安全漏洞普遍,互动性与适龄性呈负相关
- 提出儿童导向AI设计指南,适合教育科技和政策制定者参考
大语言模型在面向儿童和青少年的应用中迅速普及,亟需重新审视现有以成人为中心的AI安全框架。本文指出当前安全评测基准在覆盖年龄特异性认知、情感和社会风险方面存在不足,涵盖0-6岁幼儿期、7-12岁童年期及13-18岁青春期。为此,我们推出SproutBench,一个包含1,283个发展学基础的对抗性提示评估套件,用于探测情绪依赖、隐私泄露及模仿危险行为等风险。对47种不同大语言模型的实证评估揭示了显著的安全缺陷,并发现安全与风险预防之间存在强相关性,且互动性与适龄性呈现明显负相关。这些发现为推进以儿童为中心的AI设计与部署提供了实践指导。
原文摘要 · Abstract (English)
The rapid proliferation of large language models (LLMs) in applications targeting children and adolescents necessitates a fundamental reassessment of prevailing AI safety frameworks, which are largely tailored to adult users and neglect the distinct developmental vulnerabilities of minors. This paper highlights key deficiencies in existing LLM safety benchmarks, including their inadequate coverage of age-specific cognitive, emotional, and social risks spanning early childhood (ages 0--6), middle childhood (7--12), and adolescence (13--18). To bridge these gaps, we introduce SproutBench, an innovative evaluation suite comprising 1,283 developmentally grounded adversarial prompts designed to probe risks such as emotional dependency, privacy violations, and imitation of hazardous behaviors. Through rigorous empirical evaluation of 47 diverse LLMs, we uncover substantial safety vulnerabilities, corroborated by robust inter-dimensional correlations (e.g., between Safety and Risk Prevention) and a notable inverse relationship between Interactivity and Age Appropriateness. These insights yield practical guidelines for advancing child-centric AI design and deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。