arXiv:2601.18630cs.AIcs.HC2026-01被引 6

用真人专家评估大模型心理支持质量,发现其理性强但共情弱。

Assessing the Quality of Mental Health Support in LLM Responses through Multi-Attribute Human Evaluation

  • 请精神科专家按6维度评分,评估9个大模型的心理对话表现。
  • 大模型信息准确安全,但情感回应不稳定,开源模型更平淡。
  • 适合关注心理AI临床落地与伦理评估的研究者和开发者。

全球心理健康危机持续加剧,治疗缺口大、专业心理咨询师短缺,使大型语言模型(LLMs)成为可扩展心理支持的潜在途径。尽管LLMs能提供可及的情感协助,其可靠性、治疗相关性及与人类标准的一致性仍难评估。本文提出一种基于人类专家的评估方法,构建包含500条真实场景心理对话的数据集,评估9种不同类型的LLMs(含闭源与开源模型)生成的回复。由两名精神科训练专家独立在5点李克特量表上对每个回复进行6维评分,涵盖认知支持与情感共鸣等维度。分析显示,LLMs在提供安全、连贯、临床恰当的信息方面表现良好,但情感一致性不稳定;闭源模型(如GPT-4o)响应更平衡,而开源模型表现出更大变异性与情感贫乏。研究揭示了持续存在的认知-情感差距,强调需建立具备故障意识、以临床为基的评估框架,优先考虑关系敏感性与信息准确性并重。文章倡导引入人类监督的平衡评估协议,推动心理导向对话AI的负责任设计与临床监管。

原文摘要 · Abstract (English)

The escalating global mental health crisis, marked by persistent treatment gaps, availability, and a shortage of qualified therapists, positions Large Language Models (LLMs) as a promising avenue for scalable support. While LLMs offer potential for accessible emotional assistance, their reliability, therapeutic relevance, and alignment with human standards remain challenging to address. This paper introduces a human-grounded evaluation methodology designed to assess LLM generated responses in therapeutic dialogue. Our approach involved curating a dataset of 500 mental health conversations from datasets with real-world scenario questions and evaluating the responses generated by nine diverse LLMs, including closed source and open source models. More specifically, these responses were evaluated by two psychiatric trained experts, who independently rated each on a 5 point Likert scale across a comprehensive 6 attribute rubric. This rubric captures Cognitive Support and Affective Resonance, providing a multidimensional perspective on therapeutic quality. Our analysis reveals that LLMs provide strong cognitive reliability by producing safe, coherent, and clinically appropriate information, but they demonstrate unstable affective alignment. Although closed source models (e.g., GPT-4o) offer balanced therapeutic responses, open source models show greater variability and emotional flatness. We reveal a persistent cognitive-affective gap and highlight the need for failure aware, clinically grounded evaluation frameworks that prioritize relational sensitivity alongside informational accuracy in mental health oriented LLMs. We advocate for balanced evaluation protocols with human in the loop that center on therapeutic sensitivity and provide a framework to guide the responsible design and clinical oversight of mental health oriented conversational AI.

心理AI大模型评估情感共情临床监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。