构建首个针对儿童内容风险的开源评测基准,评估大模型拒答不当请求的能力。
MinorBench: A hand-built benchmark for content-based risks for children
- 基于真实校园场景设计风险分类体系,聚焦儿童专属内容风险。
- 测试6个主流大模型在不同提示下的安全响应差异,发现合规性波动显著。
- 适合教育AI安全研究者与产品开发者参考,推动儿童友好型模型建设。
大型语言模型正快速进入儿童生活——通过家长使用、学校引入和同龄人传播,但现有AI伦理与安全研究未能充分应对未成年人特有的内容风险。本文通过一所中学部署的基于LLM的聊天机器人真实案例,揭示学生如何使用甚至滥用该系统。基于此,我们提出针对未成年人的内容风险新分类体系,并推出MinorBench——一个开源基准,用于评估大模型对儿童不当或不安全提问的拒绝能力。我们在不同系统提示下评估了六个主流大模型,结果表明其在儿童安全合规性方面存在显著差异。研究为构建更稳健、以儿童为中心的安全机制提供了实践指导,凸显了为年轻用户量身定制AI防护系统的紧迫性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are rapidly entering children's lives - through parent-driven adoption, schools, and peer networks - yet current AI ethics and safety research do not adequately address content-related risks specific to minors. In this paper, we highlight these gaps with a real-world case study of an LLM-based chatbot deployed in a middle school setting, revealing how students used and sometimes misused the system. Building on these findings, we propose a new taxonomy of content-based risks for minors and introduce MinorBench, an open-source benchmark designed to evaluate LLMs on their ability to refuse unsafe or inappropriate queries from children. We evaluate six prominent LLMs under different system prompts, demonstrating substantial variability in their child-safety compliance. Our results inform practical steps for more robust, child-focused safety mechanisms and underscore the urgency of tailoring AI systems to safeguard young users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。