测试大模型能否像专业心理咨询师一样进行认知行为疗法。
Assessing the Effectiveness of LLMs in Delivering Cognitive Behavioral Therapy
- 用真实治疗对话数据,对比生成和检索增强两种方法
- 大模型能模仿疗法对话,但共情和一致性不足
- 适合关注AI心理辅助效果的研究者与从业者
随着全球心理健康问题日益严重,对可及且可扩展的治疗方案需求不断上升。许多用户已开始向大型语言模型(LLMs)寻求支持,尽管这些模型尚未经过心理咨询场景的验证。本文评估了LLMs在模拟认知行为疗法(CBT)专业治疗师方面的能力。基于匿名化、转录的治疗师与客户角色扮演会话数据,我们比较了两种方法:(1) 仅生成式方法;(2) 使用CBT指南进行检索增强生成(RAG)。通过自然语言生成(NLG)指标、自然语言推理(NLI)及技能评估自动化评分,评估了专有与开源模型在语言质量、语义连贯性和治疗契合度方面的表现。结果表明,虽然LLMs能生成类似CBT的对话,但在传递共情和保持一致性方面存在明显局限。
原文摘要 · Abstract (English)
As mental health issues continue to rise globally, there is an increasing demand for accessible and scalable therapeutic solutions. Many individuals currently seek support from Large Language Models (LLMs), even though these models have not been validated for use in counseling services. In this paper, we evaluate LLMs' ability to emulate professional therapists practicing Cognitive Behavioral Therapy (CBT). Using anonymized, transcribed role-play sessions between licensed therapists and clients, we compare two approaches: (1) a generation-only method and (2) a Retrieval-Augmented Generation (RAG) approach using CBT guidelines. We evaluate both proprietary and open-source models for linguistic quality, semantic coherence, and therapeutic fidelity using standard natural language generation (NLG) metrics, natural language inference (NLI), and automated scoring for skills assessment. Our results indicate that while LLMs can generate CBT-like dialogues, they are limited in their ability to convey empathy and maintain consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。