让AI理解对话中情绪的多模态成因,生成直观解释。
Generative Emotion Cause Explanation in Multimodal Conversations
- 结合文本与视频,用大模型分析面部表情捕捉情绪触发点。
- 在自建数据集ECEM上超越多个基线模型,准确率显著提升。
- 适合研究情感计算、人机交互与多模态理解的学者使用。
多模态对话是人类交流的重要形式,蕴含丰富的感情内容,探究其中情绪成因具有重要意义。然而,现有研究通常仅在单一文本模态下通过话语选择定位因果语句,存在粒度粗、解释不细致、难以识别多模态情绪诱因等问题。为此,我们提出新任务——多模态对话中的情绪原因解释(MECEC),旨在基于对话的多模态上下文生成清晰直观的情绪触发原因摘要。为支持该任务,我们基于MELD数据集构建了新数据集ECEM,融合视频片段与角色情绪详细说明,以探索多模态对话中情绪表达的因果因素。进一步提出FAME-Net方法,利用大语言模型分析视觉信息,精准解读视频中面部表情所传递的情绪。通过捕捉面部情绪的传染效应,FAME-Net有效识别出对话者的情绪成因。在新构建数据集上的实验表明,FAME-Net优于多个先进基线模型。代码与数据集已公开于https://github.com/3222345200/FAME-Net。
原文摘要 · Abstract (English)
Multimodal conversation, a crucial form of human communication, carries rich emotional content, making the exploration of the causes of emotions within it a research endeavor of significant importance. However, existing research on the causes of emotions typically employs an utterance selection method within a single textual modality to locate causal utterances. This approach remains limited to coarse-grained assessments, lacks nuanced explanations of emotional causation, and demonstrates inadequate capability in identifying multimodal emotional triggers. Therefore, we introduce a task-\textbf{Multimodal Emotion Cause Explanation in Conversation (MECEC)}. This task aims to generate a summary based on the multimodal context of conversations, clearly and intuitively describing the reasons that trigger a given emotion. To adapt to this task, we develop a new dataset (ECEM) based on the MELD dataset. ECEM combines video clips with detailed explanations of character emotions, helping to explore the causal factors behind emotional expression in multimodal conversations. A novel approach, FAME-Net, is further proposed, that harnesses the power of Large Language Models (LLMs) to analyze visual data and accurately interpret the emotions conveyed through facial expressions in videos. By exploiting the contagion effect of facial emotions, FAME-Net effectively captures the emotional causes of individuals engaged in conversations. Our experimental results on the newly constructed dataset show that FAME-Net outperforms several excellent baselines. Code and dataset are available at https://github.com/3222345200/FAME-Net.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。