构建首个多模态对话情感因果三元组数据集并提出新模型提升抽取效果
M3HG: Multimodal, Multi-scale, and Multi-type Node Heterogeneous Graph for Emotion Cause Triplet Extraction in Conversations
- 设计多模态异构图捕捉话语内外上下文关系
- 在989段对话上实现超越现有方法的三元组抽取性能
- 适合关注对话情感分析与多模态信息融合的研究者
多模态对话中的情感因果三元组抽取(MECTEC)在社交媒体分析中日益重要,旨在同时提取情感语句、原因语句及情感类别。然而,相关数据集稀缺,仅有一个公开数据集且对话场景高度同质,制约了模型发展。为此,我们提出MECAD,首个多模态、多场景的MECTEC数据集,包含来自56部电视剧的989段对话,涵盖广泛对话情境。现有MECTEC方法未能显式建模情感与因果上下文,且忽略不同层次语义信息的融合,导致性能下降。本文提出M3HG,通过多模态异构图显式捕捉情感与因果上下文,并在句间与句内层面有效融合上下文信息。大量实验表明,M3HG优于现有最先进方法。代码与数据集已开源于https://github.com/redifinition/M3HG。
原文摘要 · Abstract (English)
Emotion Cause Triplet Extraction in Multimodal Conversations (MECTEC) has recently gained significant attention in social media analysis, aiming to extract emotion utterances, cause utterances, and emotion categories simultaneously. However, the scarcity of related datasets, with only one published dataset featuring highly uniform dialogue scenarios, hinders model development in this field. To address this, we introduce MECAD, the first multimodal, multi-scenario MECTEC dataset, comprising 989 conversations from 56 TV series spanning a wide range of dialogue contexts. In addition, existing MECTEC methods fail to explicitly model emotional and causal contexts and neglect the fusion of semantic information at different levels, leading to performance degradation. In this paper, we propose M3HG, a novel model that explicitly captures emotional and causal contexts and effectively fuses contextual information at both inter- and intra-utterance levels via a multimodal heterogeneous graph. Extensive experiments demonstrate the effectiveness of M3HG compared with existing state-of-the-art methods. The codes and dataset are available at https://github.com/redifinition/M3HG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。