通过融合语音情感特征,提升长对话中情绪因果推理能力。
Decoding the Flow: CauseMotion for Emotional Causality Analysis in Long-form Conversations
- 结合语音情绪、强度与语速,增强文本语义表示
- 在70+轮对话中实现8.7%的因果准确率提升
- 适合需要深度情绪分析的对话系统研究者
长序列因果推理旨在揭示长时间序列数据中的因果关系,但受复杂依赖和因果链验证困难的制约。为克服大模型(如GPT-4)在长对话中捕捉复杂情绪因果关系的局限,本文提出CauseMotion框架,基于检索增强生成(RAG)与多模态融合。不同于仅依赖文本的方法,CauseMotion引入音频提取的声学特征——语音情绪、情感强度与语速,丰富文本表征。通过RAG与滑动窗口机制,有效检索并利用上下文相关对话片段,实现跨多轮对话的复杂情绪因果链推断。为此,我们构建首个专注长序列情绪因果推理的基准数据集,包含超过70轮对话。实验表明,该RAG增强的多模态融合方法显著提升大模型的情绪理解深度与因果推理能力。集成GLM-4的CauseMotion相较原模型提升8.7%因果准确率,超越GPT-4o 1.2%。在公开数据集DiaASQ上,CauseMotion-GLM-4在准确率、F1及因果推理准确率上均达当前最优表现。
原文摘要 · Abstract (English)
Long-sequence causal reasoning seeks to uncover causal relationships within extended time series data but is hindered by complex dependencies and the challenges of validating causal links. To address the limitations of large-scale language models (e.g., GPT-4) in capturing intricate emotional causality within extended dialogues, we propose CauseMotion, a long-sequence emotional causal reasoning framework grounded in Retrieval-Augmented Generation (RAG) and multimodal fusion. Unlike conventional methods relying only on textual information, CauseMotion enriches semantic representations by incorporating audio-derived features-vocal emotion, emotional intensity, and speech rate-into textual modalities. By integrating RAG with a sliding window mechanism, it effectively retrieves and leverages contextually relevant dialogue segments, thus enabling the inference of complex emotional causal chains spanning multiple conversational turns. To evaluate its effectiveness, we constructed the first benchmark dataset dedicated to long-sequence emotional causal reasoning, featuring dialogues with over 70 turns. Experimental results demonstrate that the proposed RAG-based multimodal integrated approach, the efficacy of substantially enhances both the depth of emotional understanding and the causal inference capabilities of large-scale language models. A GLM-4 integrated with CauseMotion achieves an 8.7% improvement in causal accuracy over the original model and surpasses GPT-4o by 1.2%. Additionally, on the publicly available DiaASQ dataset, CauseMotion-GLM-4 achieves state-of-the-art results in accuracy, F1 score, and causal reasoning accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。