arXiv:2604.17730cs.CLcs.AI2026-04ACL被引 2

提出新评估框架,检测大模型心理咨询中的角色性安全风险。

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

论文配图:MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
图 1 · 摘自论文原文
  • 基于角色意识的分类体系,识别模型在对话中扮演的有害角色。
  • 通过对抗性多轮交互发现累积性安全漏洞,覆盖率达传统方法3倍以上。
  • 适合研究者、AI医疗团队用于诊断模型潜在心理伤害机制。

大型语言模型在心理健康咨询中的应用日益广泛,但其安全性评估仍面临挑战,因临床危害具有交互性和情境依赖性。现有评估框架多基于孤立响应与粗粒度分类,难以揭示危害在多轮对话中的演变与积累。本文提出R-MHSafe:一种角色感知的心理健康安全分类体系,从AI咨询师可能扮演的施害者、挑动者、促成者或助燃者等角色出发,结合临床危害类别进行刻画。进而构建MHSafeEval——一个闭环、基于智能体的评估框架,将安全评估转化为由角色感知引导的对抗性多轮交互中的危害轨迹发现。利用该框架对多个先进大模型进行大规模评估,结果显示传统静态基准系统性遗漏了大量角色相关且累积性的安全缺陷;本方法显著提升了故障模式覆盖率与诊断精细度。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly explored as scalable tools for mental health counseling, yet evaluating their safety remains challenging due to the interactional and context-dependent nature of clinical harm. Existing evaluation frameworks predominantly assess isolated responses using coarse-grained taxonomies or static datasets, limiting their ability to diagnose how harms emerge and accumulate over multi-turn counseling interactions. In this work, we introduce R-MHSafe, a role-aware mental health safety taxonomy that characterizes clinically significant harm in terms of the interactional roles an AI counselor adopts, including perpetrator, instigator, facilitator, or enabler, combined with clinically grounded harm categories. Then, we propose MHSafeEval, a closed-loop, agent-based evaluation framework that formulates safety assessment as trajectory-level discovery of harm through adversarial multi-turn interactions, guided by role-aware modeling. Using R-MHSafe and MHSafeEval, we conduct a large-scale evaluation across state-of-the-art LLMs. Our results reveal substantial role-dependent and cumulative safety failures that are systematically missed by existing static benchmarks, and show that our framework significantly improves failure-mode coverage and diagnostic granularity.

心理健康安全评估大模型角色建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。