arXiv:2607.22692cs.AI2026-07

为生成式AI心理支持设计多轮安全架构,提升风险响应能力

Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture

  • 构建可适配多种模型的安全治理框架,融合风险识别与应答策略
  • 风险检测敏感度达0.92,特异性0.85,临床推荐转介率提升超25个百分点
  • 适用于医疗场景的AI对话系统,尤其适合需长期陪伴的心理支持应用

大型语言模型在情感支持中日益普及,但缺乏对动态心理风险的安全治理机制。现有方法多聚焦于风险识别,很少关注对话中风险演化时的应对策略。本文提出一种模型无关的安全治理架构,结合上下文风险检测、基于推理的验证和协议引导的应答生成,用于多轮心理支持对话。基于真实心理健康叙事生成的合成对话进行评估,测试模型包括GPT-5-chat和Qwen3.5-27B,结果显示风险检测表现优异(特异性:0.85(95%CI: 0.78;0.91),敏感度:0.92(95%CI: 0.88;0.95)),同时使临床医生偏好转介响应率提升25.6–59.2个百分点,且维持良好关系联结。该架构在对话长度变化下性能稳定,并在专有与开源模型间具有泛化能力。研究证明,基于临床实践的安全治理不仅能识别风险,还能优化模型对动态心理危机的应对方式,为多模型部署提供可扩展的安全框架。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existing safety approaches primarily detect risk but rarely shape how models respond as conversational risk unfolds. We developed a model-agnostic safety governance architecture that combines contextual risk detection, reasoning-based verification, and protocol-guided response generation for multi-turn mental health interactions. Synthetic conversations grounded in real-world mental health narratives were used to evaluate the architecture's performance, tested with GPT-5-chat and Qwen3.5-27B, achieving high risk detection performance (specificity: 0.85 (95\%CI: 0.78;0.91), sensitivity: 0.92 (95\%CI: 0.88;0.95)) and increasing clinician-preferred escalation responses by 25.6--59.2pp while preserving rapport and connection. Performance remained stable across conversation length and generalized across both proprietary and open-source models. These findings demonstrate that clinically-grounded safety governance can extend beyond risk detection to improve how LLMs manage evolving mental health risk, providing a scalable framework for safer deployment across models.

心理支持安全架构大模型治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。