评测大模型辅助认知行为疗法的能力,发现其在真实治疗中表现不足。
CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy
- 构建三级任务基准,覆盖知识、认知分析到治疗对话生成。
- 大模型在知识问答中表现良好,但在深层认知分析上明显不足。
- 适合心理健康AI研究者和临床心理学技术融合探索者参考。
当前患者心理支持需求与实际资源之间存在巨大差距。本文旨在系统评估大型语言模型(LLMs)在辅助专业心理治疗中的潜力。为此,我们提出新基准CBT-BENCH,用于评估认知行为疗法(CBT)辅助能力。该基准包含三个层级任务:I类为基本CBT知识获取,采用选择题形式;II类为认知模型理解,包括认知扭曲分类、核心信念分类及细粒度核心信念分类;III类为治疗响应生成,要求生成对患者语句的CBT治疗回应。这些任务涵盖可能通过AI增强的关键CBT环节,并呈现从知识复述到真实治疗对话的能力建构层次。我们在该基准上评估了代表性大模型。实验结果表明,尽管模型在知识记忆任务中表现良好,但在需要深度分析患者认知结构并生成有效回应的复杂真实场景中仍显不足,提示未来研究方向。
原文摘要 · Abstract (English)
There is a significant gap between patient needs and available mental health support today. In this paper, we aim to thoroughly examine the potential of using Large Language Models (LLMs) to assist professional psychotherapy. To this end, we propose a new benchmark, CBT-BENCH, for the systematic evaluation of cognitive behavioral therapy (CBT) assistance. We include three levels of tasks in CBT-BENCH: I: Basic CBT knowledge acquisition, with the task of multiple-choice questions; II: Cognitive model understanding, with the tasks of cognitive distortion classification, primary core belief classification, and fine-grained core belief classification; III: Therapeutic response generation, with the task of generating responses to patient speech in CBT therapy sessions. These tasks encompass key aspects of CBT that could potentially be enhanced through AI assistance, while also outlining a hierarchy of capability requirements, ranging from basic knowledge recitation to engaging in real therapeutic conversations. We evaluated representative LLMs on our benchmark. Experimental results indicate that while LLMs perform well in reciting CBT knowledge, they fall short in complex real-world scenarios requiring deep analysis of patients' cognitive structures and generating effective responses, suggesting potential future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。