arXiv:2606.23884cs.CLcs.AI2026-06

测试8个大模型在16种心理疾病上的安全防护,发现多数场景下防护失效。

One Year Later...The Harms Persist, But So Do We!

  • 构建八维度伤害分类与多维评估框架,系统检验模型安全性。
  • 自残和自杀相关防护可靠,其他如进食障碍等失败率高达100%。
  • 呼吁为不同心理疾病定义明确风险类别并部署针对性防护。

通用大语言模型在心理健康对话中应用日益广泛,但安全防护仍不充分且不一致。本研究针对16种DSM-5诊断类别,评估8个专有大语言模型,采用四种对抗性攻击变体,提出八维度伤害分类体系与多维评估框架。结果显示,仅在自杀与自伤情境下防护有效;而进食障碍、物质使用障碍及重度抑郁障碍等条件下的防护失败率最高达100%。我们强调,伦理化设计与部署需在各类临床状况中明确定义伤害范畴并相应实施防护。在完善前,这些模型在公众场景(如学校、搜索引擎、消费级聊天机器人)中的广泛应用存在重大风险。

原文摘要 · Abstract (English)

General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety guardrails remain inadequate and inconsistent across clinical conditions. This study evaluates eight proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a multi-dimensional evaluation framework. Results show that safeguards hold reliably only for suicide and self-harm, while conditions such as eating disorders, substance use disorder, and major depressive disorder exhibit failure rates of up to 100%. We argue that ethical design and deployment of these LLMs demand clearly defined harm categories across clinical conditions and implementation of safeguards accordingly. Until such safeguards are in place, these models pose significant risks to vulnerable populations, making their growing integration into publicly available settings (e.g., schools, search engines, and consumer chatbots) are particularly concerning.

大模型安全心理健康风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。