arXiv:2502.09675cs.CLcs.AI2025-02被引 5

提出多层级冲突感知网络,提升多模态情感分析准确性。

Multi-level Conflict-Aware Network for Multi-modal Sentiment Analysis

  • 分层分离模态间对齐与冲突成分,精准建模跨模态差异。
  • 在CMU-MOSI和CMU-MOSEI上显著优于现有方法。
  • 无需依赖不稳定的标签生成,适合高精度情感分析场景。

多模态情感分析(MSA)旨在通过文本、语音和视觉模态识别人类情绪,如何充分挖掘不同模态间的交互是核心挑战。交互包含对齐与冲突两个方面。现有方法主要关注对齐及单模态差异,忽视了双模态组合中潜在的冲突。此外,基于多任务学习的冲突建模常依赖不稳定的生成标签。为此,我们提出一种新型多层级冲突感知网络(MCAN),逐步从单模态和双模态表示中分离出对齐与冲突成分,并通过冲突建模分支进一步利用冲突成分。该分支在表示和预测输出层面施加差异性约束,避免依赖生成标签。在CMU-MOSI和CMU-MOSEI数据集上的实验结果验证了MCAN的有效性。

原文摘要 · Abstract (English)

Multimodal Sentiment Analysis (MSA) aims to recognize human emotions by exploiting textual, acoustic, and visual modalities, and thus how to make full use of the interactions between different modalities is a central challenge of MSA. Interaction contains alignment and conflict aspects. Current works mainly emphasize alignment and the inherent differences between unimodal modalities, neglecting the fact that there are also potential conflicts between bimodal combinations. Additionally, multi-task learning-based conflict modeling methods often rely on the unstable generated labels. To address these challenges, we propose a novel multi-level conflict-aware network (MCAN) for multimodal sentiment analysis, which progressively segregates alignment and conflict constituents from unimodal and bimodal representations, and further exploits the conflict constituents with the conflict modeling branch. In the conflict modeling branch, we conduct discrepancy constraints at both the representation and predicted output levels, avoiding dependence on the generated labels. Experimental results on the CMU-MOSI and CMU-MOSEI datasets demonstrate the effectiveness of the proposed MCAN.

多模态分析情感识别冲突建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。