arXiv:2607.24754cs.HCcs.AI2026-07

构建统一评估框架,让心理健康大模型评测更可复现、可比较。

CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs

论文配图:CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs
图 1 · 摘自论文原文
  • 提出CARE-MH统一评估框架,解决评测设计不一致问题。
  • 发现模型稳定性是复现性的关键,不同评分标准导致评测结果差异大。
  • 适合研究者和开发者用于规范心理健康大模型评估流程。

大语言模型在心理健康支持中应用日益广泛,亟需可靠评估其安全性、共情能力与治疗适宜性。然而,现有心理健康评测基准因评价设计与指标定义不一致,难以复现与比较。本文提出CARE-MH,一个统一的、可复现且可比的心理健康大模型评估框架。基于该框架,我们复现并分析了当前最先进的评测,发现复现性高度依赖模型稳定性,而跨基准分歧主要源于指标定义差异。研究强调未来心理健康大模型评测需采用标准化配置与共享指标定义。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic appropriateness. However, existing mental health benchmarks are difficult to reproduce and compare due to inconsistent evaluation designs and metric definitions. We present CARE-MH, a unified framework for comparable and reproducible evaluation of mental health LLMs. Using CARE-MH, we reproduce and analyze state-of-the-art benchmarks, revealing that reproducibility depends strongly on model stability and that cross-benchmark disagreement primarily arises from differences in metric definitions. Our findings highlight the need for standardized evaluation configurations and shared metric definitions for future mental health LLM benchmarks.

大模型评测心理健康可复现性评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。