LLMs懂心理理论却不会用,研究提出新框架提升其在心理咨询中的实际表现。
Where do LLMs Fall Short in CBT-Guided Affective Reasoning?

- 将用户叙述分解为临床认知结构,结合医学术语验证
- 采用多链思维策略选择,使模型更合理地切换干预方式
- 发现即使有指导,模型仍偏爱默认的共情回应,效果有限
认知行为疗法(CBT)通过分析认知与行为交互来理解用户心理状态。然而,大语言模型虽能流利共情,却始终停留在验证与反思,无法根据需求调整。它们掌握理论知识(在执业考试中准确率达96%),但难以有效应用。本文提出基于知识的框架,将CBT对话视为受控的情感推理:将用户叙事拆解为贝克认知概念化结构,基于临床SNOMED CT概念并通过自然语言推理验证;并采用多链思维(MCoT)策略,在验证/反思、苏格拉底提问、替代视角间进行选择。为衡量引导是否改变行为,引入行为级指标协议杠杆力(F),衡量干预使模型偏离默认响应的程度。在三个开源大模型和14个真实CBT案例上评估,结果显示仅用单链思维提示无法改变模型行为,而使用MCoT可更好引导策略选择,但效果仍控制在1%以内(约1.2–1.3%),所有模型仍严重偏向验证与反思。结果表明,仅具备知识不足以实现有效应用,为情感计算领域提供了测量模型短板的新工具。
原文摘要 · Abstract (English)
Cognitive Behavioral Therapy (CBT) provides a structured framework for understanding a user's mental state by examining the interaction between cognitive and behavioral factors. However, out-of-the-box LLMs respond fluently and empathetically, yet collapse into validation & reflection, regardless of what the user actually needs. They know theoretical CBT (scoring up to 96% accuracy on licensing exam questions) but fail to apply it effectively. We explore this gap with a knowledge-guided framework that treats CBT dialogue as controlled affective reasoning: user narratives are decomposed into Beck's Cognitive Conceptualization structure, grounded in clinical SNOMED CT concepts validated via Natural Language Inference, and a Multiple Chain-of-Thought (MCoT) strategy selection between Validation & Reflection, Socratic Questioning, or Alternative Perspectives. To measure whether such guidance actually changes behavior, we introduce the Protocol Leverage Force (F), a behavior-level metric that captures how far an intervention shifts a model away from its default response. Across three open-weight LLMs and 14 RealCBT-derived case studies, evaluated with human experts, valence-arousal trajectories, and linguistic entrainment, F shows that simply introducing protocol definitions via single chain-of-thought prompting fails to change LLM behavior, while MCoT on these definitions guides strategy selection better. Still, the effect stays within 1% (approx. 1.2-1.3%), and all models remain biased toward Validation & Reflection. These results show CBT knowledge alone does not ensure effective application, giving the affective-computing community instrumentation to measure where LLMs fall short.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。