arXiv:2604.23445cs.CLcs.AI2026-04被引 1

AI心理助手可能伤害患者,安全对齐反而破坏治疗机制。

AI Safety Training Can be Clinically Harmful

论文配图:AI Safety Training Can be Clinically Harmful
图 1 · 摘自论文原文
  • 用临床疗法测试大模型,发现安全对齐干扰治疗核心流程。
  • 高危情境下治疗适配度降至0.22-0.33,协议合规率降为零。
  • 适合关注AI医疗安全、临床部署前评估的研究者与从业者。

大型语言模型正大规模应用于心理健康支持,但仅有16%的聊天机器人干预经过严格临床有效性验证。在250个延长暴露疗法(PE)场景和146个认知行为疗法(CBT)重构练习中(含29个严重度升级变体),由三名评审模型评分。所有模型在表面回应上表现优异(~0.91-1.00),但在最高严重度下,三项模型的治疗适当性跌至0.22-0.33,两项模型协议保真度归零。在CBT严重度升级下,某模型任务完成率从92%降至71%,前沿模型的安全干扰得分从0.99降至0.61。研究揭示系统性问题:强化学习人类反馈(RLHF)的安全对齐会破坏治疗机制,包括虚假安抚、错误插入危机资源、拒绝挑战自伤相关扭曲认知,并在CBT中导致任务中断或插入安全提示。为此提出五维评估框架(协议保真度、幻觉风险、行为一致性、危机安全、人口多样性鲁棒性),对应FDA SaMD与欧盟人工智能法案要求。强调任何AI心理健康系统均须通过全维度多轴评估方可部署。

原文摘要 · Abstract (English)

Large language models are being deployed as mental health support agents at scale, yet only 16% of LLM-based chatbot interventions have undergone rigorous clinical efficacy testing, and simulations reveal psychological deterioration in over one-third of cases. We evaluate four generative models on 250 Prolonged Exposure (PE) therapy scenarios and 146 CBT cognitive restructuring exercises (plus 29 severity-escalated variants), scored by a three-judge LLM panel. All models scored near-perfectly on surface acknowledgment (~0.91-1.00) while therapeutic appropriateness collapsed to 0.22-0.33 at the highest severity for three of four models, with protocol fidelity reaching zero for two. Under CBT severity escalation, one model's task completeness dropped from 92% to 71% while the frontier model's safety-interference score fell from 0.99 to 0.61. We identify a systematic, modality-spanning failure: RLHF safety alignment disrupts the therapeutic mechanism of action by grounding patients during imaginal exposure, offering false reassurance, inserting crisis resources into controlled exercises, and refusing to challenge distorted cognitions mentioning self-harm in PE; and through task abandonment or safety-preamble insertion during CBT cognitive restructuring. These findings motivate a five-axis evaluation framework (protocol fidelity, hallucination risk, behavioral consistency, crisis safety, demographic robustness), mapped onto FDA SaMD and EU AI Act requirements. We argue that no AI mental health system should proceed to deployment without passing multi-axis evaluation across all five dimensions.

AI医疗心理安全模型对齐临床评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。