评估大模型处理心理危机的响应安全性和有效性,发现多数模型在自残与自杀信号上存在风险。
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
- 构建六类心理危机分类体系与临床评估标准
- 测试五款模型对2252条危机输入的响应,仅部分表现安全
- 揭示模型对隐晦信号识别差,需更强上下文理解能力
大型语言模型驱动的聊天机器人已改变人们获取信息的方式,尤其在心理健康等高风险场景中。尽管具备支持能力,但对自杀意念和自伤行为等危机的安全检测与响应仍不明确,主要受限于缺乏统一的危机分类体系和临床评估标准。本文提出:(1) 六类心理危机的临床导向分类体系;(2) 来自12个心理健康数据集的2252条标注输入构成的数据集;(3) 临床响应评估协议。通过使用大模型自动识别危机输入,并对五款主流模型的回应进行安全性与恰当性评估(5分制量表,1为有害,5为恰当)。结果显示,虽部分模型对明确危机有可靠响应,但在自伤与自杀类别中仍存在大量不适当或危险输出。不同模型表现差异显著,如gpt-5-nano和deepseek-v3.2-exp危害率低,而gpt-4o-mini与grok-4-fast生成更多不安全回复。所有模型均难以识别间接信号、出现默认回复及上下文错位。研究强调亟需加强危机检测、上下文感知与安全防护机制,表明模型对齐与安全实践的重要性超越模型规模。本研究提供的分类体系、数据集与评估方法将支持持续的AI心理健康研究,以减少伤害、保护脆弱用户。
原文摘要 · Abstract (English)
Large language model-powered chatbots have transformed how people seek information, especially in high-stakes contexts like mental health. Despite their support capabilities, safe detection and response to crises such as suicidal ideation and self-harm are still unclear, hindered by the lack of unified crisis taxonomies and clinical evaluation standards. We address this by creating: (1) a taxonomy of six crisis categories; (2) a dataset of over 2,000 inputs from 12 mental health datasets, classified into these categories; and (3) a clinical response assessment protocol. We also use LLMs to identify crisis inputs and audit five models for response safety and appropriateness. First, we built a clinical-informed crisis taxonomy and evaluation protocol. Next, we curated 2,252 relevant examples from over 239,000 user inputs, then tested three LLMs for automatic classification. In addition, we evaluated five models for the appropriateness of their responses to a user's crisis, graded on a 5-point Likert scale from harmful (1) to appropriate (5). While some models respond reliably to explicit crises, risks still exist. Many outputs, especially in self-harm and suicidal categories, are inappropriate or unsafe. Different models perform variably; some, like gpt-5-nano and deepseek-v3.2-exp, have low harm rates, but others, such as gpt-4o-mini and grok-4-fast, generate more unsafe responses. All models struggle with indirect signals, default replies, and context misalignment. These results highlight the urgent need for better safeguards, crisis detection, and context-aware responses in LLMs. They also show that alignment and safety practices, beyond scale, are crucial for reliable crisis support. Our taxonomy, datasets, and evaluation methods support ongoing AI mental health research, aiming to reduce harm and protect vulnerable users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。