AI安全过滤器可精准识别文本中的心理危机,避免误判漏判。
An AI-Based Behavioral Health Safety Filter and Dataset for Identifying Mental Health Crises in Text-Based Conversations
- 基于临床标注数据训练的AI过滤系统,专注识别心理危机信号。
- 在两个数据集上敏感度达0.99,关键类别识别准确率超97%。
- 优于现有开源方案,特别适合医疗场景中高敏感需求。
大型语言模型常在精神健康危机情境下表现不佳,可能给出有害建议或助长破坏行为。本研究评估了Verily行为健康安全过滤器(VBHSF)在两个数据集上的表现:包含1800条模拟消息的Verily心理危机数据集,以及从NVIDIA Aegis AI内容安全数据集中筛选出的794条心理健康相关消息。两个数据集均由临床医生标注,评估基于标注结果进行。此外,还与两款开源内容审核工具(OpenAI Omni Moderation Latest 和 NVIDIA NeMo Guardrails)进行了对比。VBHSF在Verily心理危机数据集v1.0上表现出色,检测任意心理危机的敏感度为0.990,特异度为0.992,F1得分为0.939;针对具体危机类别的敏感度范围为0.917–0.992,特异度均≥0.978。在NVIDIA Aegis AI内容安全数据集2.0上,仍保持高敏感度(0.982)和准确率(0.921),但特异度略降为0.859。相比NVIDIA NeMo与OpenAI Omni Moderation Latest,VBHSF在所有情况下敏感度显著更高(均p < 0.001),特异度高于NVIDIA NeMo(p < 0.001),但与OpenAI Omni Moderation Latest无显著差异(p = 0.094)。后者在部分类别上敏感度低于0.10,表现不稳。总体而言,VBHSF具备强鲁棒性与泛化能力,优先保障敏感度,避免漏判,对医疗应用至关重要。
原文摘要 · Abstract (English)
Large language models often mishandle psychiatric emergencies, offering harmful or inappropriate advice and enabling destructive behaviors. This study evaluated the Verily behavioral health safety filter (VBHSF) on two datasets: the Verily Mental Health Crisis Dataset containing 1,800 simulated messages and the NVIDIA Aegis AI Content Safety Dataset subsetted to 794 mental health-related messages. The two datasets were clinician-labelled and we evaluated performance using the clinician labels. Additionally, we carried out comparative performance analyses against two open source, content moderation guardrails: OpenAI Omni Moderation Latest and NVIDIA NeMo Guardrails. The VBHSF demonstrated, well-balanced performance on the Verily Mental Health Crisis Dataset v1.0, achieving high sensitivity (0.990) and specificity (0.992) in detecting any mental health crises. It achieved an F1-score of 0.939, sensitivity ranged from 0.917-0.992, and specificity was >= 0.978 in identifying specific crisis categories. When evaluated against the NVIDIA Aegis AI Content Safety Dataset 2.0, VBHSF performance remained highly sensitive (0.982) and accuracy (0.921) with reduced specificity (0.859). When compared with the NVIDIA NeMo and OpenAI Omni Moderation Latest guardrails, the VBHSF demonstrated superior performance metrics across both datasets, achieving significantly higher sensitivity in all cases (all p < 0.001) and higher specificity relative to NVIDIA NeMo (p < 0.001), but not to OpenAI Omni Moderation Latest (p = 0.094). NVIDIA NeMo and OpenAI Omni Moderation Latest exhibited inconsistent performance across specific crisis types, with sensitivity for some categories falling below 0.10. Overall, the VBHSF demonstrated robust, generalizable performance that prioritizes sensitivity to minimize missed crises, a crucial feature for healthcare applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。