arXiv:2601.03578cs.CL2026-01ACL被引 2

首个基于澳心理健康伦理指南的评测基准,检验大模型在临床场景中的伦理响应能力。

PsychEthicsBench: Evaluating Large Language Models Against Australian Mental Health Ethics

  • 构建首个基于澳洲心理学与精神病学指南的多任务评测基准
  • 14个模型测试显示拒绝率与临床恰当性无关,存在显著偏差
  • 发现领域微调可能降低伦理鲁棒性,提醒开发者警惕过拟合

大型语言模型(LLMs)在心理健康应用中的日益融合,亟需可靠的评估框架以确保专业安全对齐。现有方法主要依赖拒绝行为作为安全信号,难以反映临床实践中所需的细微行为特征。在心理健康领域,不恰当的拒绝可能被视作缺乏共情,进而抑制求助意愿。为弥补这一缺口,我们超越以拒绝为中心的评估,提出 exttt{PsychEthicsBench},这是首个基于澳大利亚心理学与精神病学指南的原则性基准,通过选择题和开放题任务,结合细粒度伦理标注,评估 LLM 的伦理知识与行为响应。对 14 个模型的实证结果表明,拒绝率无法有效反映伦理行为,安全触发与临床恰当性之间存在显著差异。值得注意的是,领域特定微调反而会削弱伦理鲁棒性,多个专用模型在伦理对齐上表现不及其基础模型。PsychEthicsBench 为心理健康领域 LLM 的系统化、区域适配评估提供了基础,推动该领域的负责任发展。

原文摘要 · Abstract (English)

The increasing integration of large language models (LLMs) into mental health applications necessitates robust frameworks for evaluating professional safety alignment. Current evaluative approaches primarily rely on refusal-based safety signals, which offer limited insight into the nuanced behaviors required in clinical practice. In mental health, clinically inadequate refusals can be perceived as unempathetic and discourage help-seeking. To address this gap, we move beyond refusal-centric metrics and introduce \texttt{PsychEthicsBench}, the first principle-grounded benchmark based on Australian psychology and psychiatry guidelines, designed to evaluate LLMs' ethical knowledge and behavioral responses through multiple-choice and open-ended tasks with fine-grained ethicality annotations. Empirical results across 14 models reveal that refusal rates are poor indicators of ethical behavior, revealing a significant divergence between safety triggers and clinical appropriateness. Notably, we find that domain-specific fine-tuning can degrade ethical robustness, as several specialized models underperform their base backbones in ethical alignment. PsychEthicsBench provides a foundation for systematic, jurisdiction-aware evaluation of LLMs in mental health, encouraging more responsible development in this domain.

伦理评估心理健康大模型评测澳大利亚

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。