arXiv:2502.06873cs.CLcs.AI2025-02NAACL被引 9

让AI通过表情和对话理解情绪,更像真人心理医生。

Multimodal Cognitive Reframing Therapy via Multi-hop Psychotherapeutic Reasoning

  • 用图像+对话构建多模态心理治疗数据集
  • 视觉线索提升AI对隐性情绪的识别能力
  • 适合研究人机共情与智能心理辅助系统

先前研究显示大语言模型(LLMs)在认知重构疗法中具有潜力,但主要聚焦于文本方法,常忽略真实治疗中非语言证据的重要性。为弥补这一差距,我们拓展了基于文本的认知重构至多模态,引入视觉线索。具体而言,提出新数据集Multi Modal-Cognitive Support Conversation (M2CoSC),将GPT-4生成的对话与反映虚拟来访者面部表情的图像配对。为更贴近真实心理治疗中由面部表情推断隐性情绪证据的过程,我们提出多跳心理治疗推理方法,显式识别并整合细微线索。在LLMs与视觉语言模型(VLMs)上的全面实验表明,使用M2CoSC数据集后,VLM作为心理治疗师的表现显著提升;且多跳推理方法使VLM能提供更深入、更具同理心的建议,优于标准提示方法。

原文摘要 · Abstract (English)

Previous research has revealed the potential of large language models (LLMs) to support cognitive reframing therapy; however, their focus was primarily on text-based methods, often overlooking the importance of non-verbal evidence crucial in real-life therapy. To alleviate this gap, we extend the textual cognitive reframing to multimodality, incorporating visual clues. Specifically, we present a new dataset called Multi Modal-Cognitive Support Conversation (M2CoSC), which pairs each GPT-4-generated dialogue with an image that reflects the virtual client's facial expressions. To better mirror real psychotherapy, where facial expressions lead to interpreting implicit emotional evidence, we propose a multi-hop psychotherapeutic reasoning approach that explicitly identifies and incorporates subtle evidence. Our comprehensive experiments with both LLMs and vision-language models (VLMs) demonstrate that the VLMs' performance as psychotherapists is significantly improved with the M2CoSC dataset. Furthermore, the multi-hop psychotherapeutic reasoning method enables VLMs to provide more thoughtful and empathetic suggestions, outperforming standard prompting methods.

心理治疗多模态大模型情感识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。