针对儿童使用大模型的安全风险,提出专属评估框架并发现显著安全漏洞。
LLM Safety for Children
- 构建基于儿童心理学的用户模型,模拟不同性格与兴趣的儿童。
- 测试6个主流大模型,发现对儿童有害的内容过滤严重不足。
- 为教育、心理辅导等儿童场景提供可落地的安全评估方案。
本文分析了大型语言模型(LLMs)在18岁以下儿童使用场景中的安全性问题。尽管LLMs已在教育、心理治疗等领域广泛应用,但针对儿童群体的内容风险评估仍存在明显空白。研究指出,儿童在认知、情感与行为上具有多样性,常被标准安全评估忽视。为此,本文基于儿童发展心理学文献,构建了反映儿童个性与兴趣差异的「儿童用户模型」,旨在填补跨领域儿童安全研究的缺口。利用该模型对六种前沿大模型进行评估,结果显示:现有模型在识别对儿童有害内容方面存在显著缺陷,尤其在涉及情感操纵、不当引导和潜在误导性信息方面表现不佳。研究强调需建立专属于儿童的评估体系,以保障其在数字环境中的安全。
原文摘要 · Abstract (English)
This paper analyzes the safety of Large Language Models (LLMs) in interactions with children below age of 18 years. Despite the transformative applications of LLMs in various aspects of children's lives such as education and therapy, there remains a significant gap in understanding and mitigating potential content harms specific to this demographic. The study acknowledges the diverse nature of children often overlooked by standard safety evaluations and proposes a comprehensive approach to evaluating LLM safety specifically for children. We list down potential risks that children may encounter when using LLM powered applications. Additionally we develop Child User Models that reflect the varied personalities and interests of children informed by literature in child care and psychology. These user models aim to bridge the existing gap in child safety literature across various fields. We utilize Child User Models to evaluate the safety of six state of the art LLMs. Our observations reveal significant safety gaps in LLMs particularly in categories harmful to children but not adults
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。