arXiv:2510.19032cs.CLcs.CY2025-10Conference of the …被引 28

构建大规模心理支持对话评估基准,检验大模型判断可靠性。

When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation

  • 整合三类真实场景数据,生成10万组对话响应对
  • 发现大模型在认知支持上可靠,但共情与安全评估偏差明显
  • 提出统计框架量化人工与模型评分一致性,适合研究者验证

评估用于心理健康的大型语言模型(LLMs)面临挑战,因治疗对话兼具情感与认知复杂性。现有基准规模小、可靠性低,多依赖合成或社交媒体数据,缺乏对自动化评判可信度的评估框架。为此,我们引入两个基准:MentalBench-100k 汇聚三个真实场景数据集的10,000轮单回合对话,每条配九个LLM生成回复,共形成100,000组响应对;MentalAlign-70k 将四个高性能LLM评判者与人类专家在70,000次评分中对比,涵盖七项属性,分为认知支持分(CSS)和情感共鸣分(ARS)。我们采用情感-认知一致性框架,通过组内相关系数(ICC)与置信区间量化模型与人类在一致性、可重复性及偏倚方面的差异。分析显示LLM评判存在系统性高估,认知类属性如引导性与信息量具强可靠性,共情评分精度下降,安全性和相关性部分不可靠。本研究建立可靠的大规模评估方法论与实证基础。

原文摘要 · Abstract (English)

Evaluating Large Language Models (LLMs) for mental health support is challenging due to the emotionally and cognitively complex nature of therapeutic dialogue. Existing benchmarks are limited in scale, reliability, often relying on synthetic or social media data, and lack frameworks to assess when automated judges can be trusted. To address the need for large-scale dialogue datasets and judge reliability assessment, we introduce two benchmarks that provide a framework for generation and evaluation. MentalBench-100k consolidates 10,000 one-turn conversations from three real scenarios datasets, each paired with nine LLM-generated responses, yielding 100,000 response pairs. MentalAlign-70k}reframes evaluation by comparing four high-performing LLM judges with human experts across 70,000 ratings on seven attributes, grouped into Cognitive Support Score (CSS) and Affective Resonance Score (ARS). We then employ the Affective Cognitive Agreement Framework, a statistical methodology using intraclass correlation coefficients (ICC) with confidence intervals to quantify agreement, consistency, and bias between LLM judges and human experts. Our analysis reveals systematic inflation by LLM judges, strong reliability for cognitive attributes such as guidance and informativeness, reduced precision for empathy, and some unreliability in safety and relevance. Our contributions establish new methodological and empirical foundations for reliable, large-scale evaluation of LLMs in mental health. We release the benchmarks and codes at: https://github.com/abeerbadawi/MentalBench/

心理健康大模型评估基准测试一致性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。