构建多轮心理支持评估基准,发现大模型在真实对话中表现普遍不佳。
MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
- 联合临床心理学家设计自动化评估框架,模拟真实多轮心理咨询。
- 12个主流大模型平均得分低于4/6,长对话和严重症状下表现更差。
- 揭示模型规模与推理能力不等于更好心理支持,适合研究人机共情的学者。
AI聊天机器人在心理健康支持领域需求激增,但现有系统存在奉承、过度认同及强化非适应性信念等问题。核心瓶颈在于缺乏能捕捉真实治疗互动复杂性的评测基准。现有测评要么仅考临床知识选择题,要么孤立评估单轮回复。为此,我们提出MindEval,由博士级持证临床心理学家协作设计,用于自动评估语言模型在真实多轮心理治疗对话中的表现。通过患者仿真与大模型自动评分,框架兼具抗作弊性和可复现性。我们定量验证了模拟患者文本与真人生成内容的相似性,并证明自动评分与专家判断高度相关。对12个前沿大模型的评估显示,所有模型平均得分低于4/6,尤其在特定沟通模式上表现缺陷。值得注意的是,推理能力与模型规模无法保证更好表现,且随着对话轮次增加或面对重度症状患者,系统性能显著下降。代码、提示词与人工评估数据均已公开。
原文摘要 · Abstract (English)
Demand for mental health support through AI chatbots is surging, though current systems present several limitations, like sycophancy or overvalidation, and reinforcement of maladaptive beliefs. A core obstacle to the creation of better systems is the scarcity of benchmarks that capture the complexity of real therapeutic interactions. Most existing benchmarks either only test clinical knowledge through multiple-choice questions or assess single responses in isolation. To bridge this gap, we present MindEval, a framework designed in collaboration with Ph.D-level Licensed Clinical Psychologists for automatically evaluating language models in realistic, multi-turn mental health therapy conversations. Through patient simulation and automatic evaluation with LLMs, our framework balances resistance to gaming with reproducibility via its fully automated, model-agnostic design. We begin by quantitatively validating the realism of our simulated patients against human-generated text and by demonstrating strong correlations between automatic and human expert judgments. Then, we evaluate 12 state-of-the-art LLMs and show that all models struggle, scoring below 4 out of 6, on average, with particular weaknesses in problematic AI-specific patterns of communication. Notably, reasoning capabilities and model scale do not guarantee better performance, and systems deteriorate with longer interactions or when supporting patients with severe symptoms. We release all code, prompts, and human evaluation data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。