评估大模型在多轮心理健康对话中的表现,发现其患者导向能力不足。
Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations
- 构建合成对话框架MedAgent生成2200+真实心理对话数据
- 顶尖推理模型在患者导向对话中平均得分仅31%
- 模型表现受用户人格影响,对话轮次越多越差
由于心理医疗服务获取受限、等待时间长以及大语言模型(LLMs)能力提升,越来越多用户转向LLMs满足心理健康需求。然而,对LLMs在多轮心理健康对话中的能力研究仍不充分。现有评估框架多关注诊断准确率和胜率,忽视了与患者个人目标、价值观和性格相匹配的沟通能力。为此,我们提出MedAgent框架,用于合成生成真实、多轮的心理健康意义建构对话,并构建了包含超过2,200个患者-模型对话的MHSD数据集。同时,我们提出了MultiSenseEval评估框架,基于以人为本的标准评估LLMs在医疗场景下的多轮对话能力。结果表明,前沿推理模型在以患者为中心的沟通中表现不佳,且在高级诊断能力上平均得分仅为31%。此外,模型表现随患者人格特征变化,对话轮次增加时性能下降。本工作提供了完整的合成数据生成框架、数据集和评估体系,用于评估LLMs在多轮心理健康对话中的表现。
原文摘要 · Abstract (English)
Limited access to mental healthcare, extended wait times, and increasing capabilities of Large Language Models (LLMs) has led individuals to turn to LLMs for fulfilling their mental health needs. However, examining the multi-turn mental health conversation capabilities of LLMs remains under-explored. Existing evaluation frameworks typically focus on diagnostic accuracy and win-rates and often overlook alignment with patient-specific goals, values, and personalities required for meaningful conversations. To address this, we introduce MedAgent, a novel framework for synthetically generating realistic, multi-turn mental health sensemaking conversations and use it to create the Mental Health Sensemaking Dialogue (MHSD) dataset, comprising over 2,200 patient-LLM conversations. Additionally, we present MultiSenseEval, a holistic framework to evaluate the multi-turn conversation abilities of LLMs in healthcare settings using human-centric criteria. Our findings reveal that frontier reasoning models yield below-par performance for patient-centric communication and struggle at advanced diagnostic capabilities with average score of 31%. Additionally, we observed variation in model performance based on patient's persona and performance drop with increasing turns in the conversation. Our work provides a comprehensive synthetic data generation framework, a dataset and evaluation framework for assessing LLMs in multi-turn mental health conversations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。