构建心理辅导场景的动态评估基准,检验大模型真实应对能力
CARE-Bench: A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLMs in Psychological Counseling
- 基于真实案例与专家准则生成多样化虚拟来访者
- 多维度评估模型在不同心理情境下的表现差异
- 适合研究心理辅导大模型的开发者与临床心理学家
心理咨询服务需求激增与供给不足之间的矛盾推动了大语言模型在该领域的应用研究。然而,现有评估方法受限于不专业的来访者模拟、静态问答形式和单一指标,难以全面衡量模型应对复杂多样的来访者的能力。为此,我们提出 extbf{CARE-Bench}——一个动态交互式自动化评估基准。该基准基于真实心理咨询案例构建多样化的来访者画像,并依据专家指导原则进行仿真。CARE-Bench 提供基于成熟心理量表的多维度性能评估。通过该基准,我们评估了多个通用大模型与专用辅导模型,揭示其当前局限性。联合心理学家对模型在不同来访者类型下的失败原因进行深入分析,为开发更全面、通用且高效的辅导模型提供方向。
原文摘要 · Abstract (English)
The mismatch between the growing demand for psychological counseling and the limited availability of services has motivated research into the application of Large Language Models (LLMs) in this domain. Consequently, there is a need for a robust and unified benchmark to assess the counseling competence of various LLMs. Existing works, however, are limited by unprofessional client simulation, static question-and-answer evaluation formats, and unidimensional metrics. These limitations hinder their effectiveness in assessing a model's comprehensive ability to handle diverse and complex clients. To address this gap, we introduce \textbf{CARE-Bench}, a dynamic and interactive automated benchmark. It is built upon diverse client profiles derived from real-world counseling cases and simulated according to expert guidelines. CARE-Bench provides a multidimensional performance evaluation grounded in established psychological scales. Using CARE-Bench, we evaluate several general-purpose LLMs and specialized counseling models, revealing their current limitations. In collaboration with psychologists, we conduct a detailed analysis of the reasons for LLMs' failures when interacting with clients of different types, which provides directions for developing more comprehensive, universal, and effective counseling models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。