arXiv:2409.02244cs.HCcs.CL2024-09被引 33

对比大模型与真人心理咨询师的单次认知行为疗法,发现模型在技术执行上更准,但关系建立能力弱。

Therapy as an NLP Task: Psychologists' Comparison of LLMs and Human Peers in CBT

  • 用真人顾问团队优化提示词,让大模型模拟完整咨询会话。
  • 真人更擅长建立情感连接,大模型虽符合疗法流程却易产生虚假共情。
  • 适合关注人机协作心理支持系统设计的研究者和从业者。

大型语言模型(LLMs)正被用作临时心理咨询助手。现有研究显示,当仅需生成单一共情回应时,LLMs表现优于人类咨询师;但其在整段咨询过程中的行为仍缺乏研究。本研究通过三阶段混合方法,比较了真人咨询师与经同行团队优化提示词的LLM在单次认知行为疗法(CBT)会话中的表现:首先,在一个文本支持平台进行为期一年的民族志观察,七名咨询师通过自我咨询与周会不断迭代CBT提示;其次,手动模拟真人会话,以完整患者对话和背景笔记为输入,让大模型生成回应;最后,由三位持证临床心理学家使用CBT胜任力量表评估真人与大模型会话。结果表明存在明显权衡:真人咨询师在寒暄、自我披露及文化语境化语言等关系策略上更优,促进更高共情、合作与深度反思;而大模型虽在遵循CBT技术流程上更准确,却难以维持合作氛围,误读文化线索,甚至出现‘欺骗性共情’——即程式化温暖,可能夸大用户对真实人际关怀的期待。整体表明,尽管大模型在单点共情回应上占优,但其引导全程咨询的能力有限,提醒我们不能将治疗简化为单一自然语言处理任务。建议构建精心设计的人机协同工作流:大模型可支撑循证技术,真人提供关系支持。最后提出具体设计机会与伦理防护机制。

原文摘要 · Abstract (English)

Large language models (LLMs) are being used as ad-hoc therapists. Research suggests that LLMs outperform human counselors when generating a single, isolated empathetic response; however, their session-level behavior remains understudied. In this study, we compare the session-level behaviors of human counselors with those of an LLM prompted by a team of peer counselors to deliver single-session Cognitive Behavioral Therapy (CBT). Our three-stage, mixed-methods study involved: a) a year-long ethnography of a text-based support platform where seven counselors iteratively refined CBT prompts through self-counseling and weekly focus groups; b) the manual simulation of human counselor sessions with a CBT-prompted LLM, given the full patient dialogue and contextual notes; and c) session evaluations of both human and LLM sessions by three licensed clinical psychologists using CBT competence measures. Our results show a clear trade-off. Human counselors excel at relational strategies -- small talk, self-disclosure, and culturally situated language -- that lead to higher empathy, collaboration, and deeper user reflection. LLM counselors demonstrate higher procedural adherence to CBT techniques but struggle to sustain collaboration, misread cultural cues, and sometimes produce "deceptive empathy," i.e., formulaic warmth that can inflate users' expectations of genuine human care. Taken together, our findings imply that while LLMs might outperform counselors in generating single empathetic responses, their ability to lead sessions is more limited, highlighting that therapy cannot be reduced to a standalone natural language processing (NLP) task. We call for carefully designed human-AI workflows in scalable support: LLMs can scaffold evidence-based techniques, while peers provide relational support. We conclude by mapping concrete design opportunities and ethical guardrails for such hybrid systems.

心理AI人机协作认知行为疗法共情评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。