arXiv:2608.07902cs.CYcs.AI2026-08

让一线儿童安全专家评估AI聊天机器人在真实危机中的表现。

Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety

论文配图:Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
图 1 · 摘自论文原文
  • 访谈19位心理社工,分析他们对聊天机器人回应的实战判断。
  • 发现现有评估忽略真实伤害,仅关注表面拒绝或恶意提示。
  • 建议将一线专家意见纳入AI安全评测体系,提升实用性。

青少年越来越多地向AI聊天机器人寻求社交与情感支持,引发对其在高风险情境下响应能力的担忧。然而,当前针对儿童安全的AI评估缺乏对青少年实际遭遇伤害的真实依据,依赖未经验证的假设(如拒绝回应即为安全),且主要聚焦于检测对抗性提示或输出中的表层危害。这些方法往往无法识别实践中可能造成伤害的回复。为更深入理解评估方法的局限性,研究团队采访了19位直接服务脆弱青少年的从业者,包括社会工作者、治疗师和心理学家,让他们基于已有实证研究中常见的青少年风险情境,反思聊天机器人的回应。从业者指出了可能导致伤害的聊天机器人行为,也识别出能真正支持青少年的关键回应方式,并讨论了聊天机器人应扮演与不应扮演的角色,提出了具体的改进建议。基于此,研究提出改进AI儿童安全评估与基础设施的具体建议,并强调应在安全工作中融入一线实践者的视角。

原文摘要 · Abstract (English)

Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations. However, existing child safety evaluations of AI lack grounding in real-world harms that youth experience, rely on unvalidated assumptions about what counts as an appropriate output (e.g., refusal), and typically focus on detecting adversarial prompts or surface-level harms in outputs only. Thus, these evaluations can fail to detect responses that pose harm to youth in practice. To better understand the limitations of current evaluation practices, we conducted interviews with 19 practitioners working directly with youth in vulnerable situations, including social workers, therapists, and psychologists, asking them to reflect on chatbots' responses to risky situations commonly faced by youth, as established in prior empirical work. Practitioners identified chatbot behaviors likely to cause harm as well as those that could meaningfully support youth in difficult moments, discussed the role that chatbots should (and should not) play in these interactions, and offered concrete recommendations for improving chatbot responses. Based on these findings, we provide recommendations for AI child safety evaluation and infrastructure, and highlight the need for incorporating practitioners' perspectives into safety work.

AI安全儿童保护评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。