arXiv:2608.25071cs.CLcs.AI2026-08

构建首个公开心理健康评估基准,助力LLM在心理支持场景的可信评测。

HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench

论文配图:HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
图 1 · 摘自论文原文
  • 基于透明规则筛选5000条医患对话,提取12.2%心理相关样本。
  • 20个主流模型测试显示排名高度一致(τ≥0.92),部分模型有明确拒绝行为。
  • 提供可复用数据集、评测流程与代码,适配开发者心理服务评估需求。

通用医疗基准日益成为衡量大语言模型医学能力的基础,但其未按临床专科细分,难以分离特定领域表现。心理健康问题关乎公共健康,数百万用户依赖LLM获取心理支持,而现有评估多为学术定制基准,难融入开发流程。我们提出HealthBench-Psych与HealthBench-Psych-Hard:通过透明的LLM评分标准筛选HealthBench中5000条医生评分对话,识别心理相关条目;经两轮盲评临床医生评审并设置已知排除对照组,最终获得610条对话(占总量12.2%)。在跨厂商三模型评委下评估20个前沿及开源模型,发现前沿模型表现统计上无显著差异,两个模型表现出可测量的拒绝行为,且评委间排名高度一致(τ≥0.92)。我们公开该子集、处理管道、模型响应、评分与分析代码,作为可复用资源。

原文摘要 · Abstract (English)

General-purpose health benchmarks increasingly anchor claims about LLM medical performance, but they are not always resolved by clinical specialty, making domain-specific performance hard to isolate. Mental health is of acute public-health concern as millions of people turn to LLMs for psychological support, and most existing evaluations are bespoke academic benchmarks that are difficult to integrate into developer workflows. We introduce HealthBench-Psych and HealthBench-Psych-Hard. We screened HealthBench's 5,000 physician-rubric conversations for mental-health relevance with a transparent LLM-applied rubric, then validated the subset through two rounds of blinded clinician review with concealed known-exclude controls, yielding 610 conversations (12.2% of the corpus). Evaluating 20 frontier and open models under a cross-vendor panel of three LLM judges, we find a statistically tied frontier cluster, measurable refusal behavior in two models, and near-identical rankings across judges ($τ\ge 0.92$). We release the subset, pipeline, model responses, grades, and analysis code as a reusable resource.

心理健康评测基准LLM评估医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。