构建心理健康大模型可信度评估基准,发现主流模型普遍存在缺陷
TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health
- 设计八维度评估框架,将心理健康领域规范转化为可量化指标
- 测试12个模型均表现不佳,连GPT-5.1也未能在所有维度达标
- 适合研究者、开发者用于系统性提升心理健康AI的安全性
尽管大型语言模型在提供可及的心理健康支持方面展现出巨大潜力,但其实际部署因领域高风险与安全敏感性而引发可信度担忧。现有通用大模型评估范式无法捕捉心理健康特定需求,亟需针对性提升其可信度。为此,我们提出TrustMH-Bench,一个系统性量化心理健康大模型可信度的综合框架。通过建立领域规范到量化指标的深度映射,该框架从可靠性、危机识别与转介、安全性、公平性、隐私性、鲁棒性、反谄媚性及伦理合规性八个核心维度进行评估。我们在六款通用大模型和六款专业心理健康模型上开展广泛实验,结果表明,所评估模型在心理健康场景中各维度表现均不理想,存在显著缺陷。值得注意的是,即使性能较强的模型(如GPT-5.1)也无法在所有维度保持高水平表现。因此,系统性提升大模型在心理健康领域的可信度已成为关键任务。相关数据与代码已公开。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) demonstrate significant potential in providing accessible mental health support, their practical deployment raises critical trustworthiness concerns due to the domains high-stakes and safety-sensitive nature. Existing evaluation paradigms for general-purpose LLMs fail to capture mental health-specific requirements, highlighting an urgent need to prioritize and enhance their trustworthiness. To address this, we propose TrustMH-Bench, a holistic framework designed to systematically quantify the trustworthiness of mental health LLMs. By establishing a deep mapping from domain-specific norms to quantitative evaluation metrics, TrustMH-Bench evaluates models across eight core pillars: Reliability, Crisis Identification and Escalation, Safety, Fairness, Privacy, Robustness, Anti-sycophancy, and Ethics. We conduct extensive experiments across six general-purpose LLMs and six specialized mental health models. Experimental results indicate that the evaluated models underperform across various trustworthiness dimensions in mental health scenarios, revealing significant deficiencies. Notably, even generally powerful models (e.g., GPT-5.1) fail to maintain consistently high performance across all dimensions. Consequently, systematically improving the trustworthiness of LLMs has become a critical task. Our data and code are released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。