用因果推理解决多模态情感分析中的偏见问题,提升分类准确率。
Multimodal Sentiment Analysis Based on Causal Reasoning
- 基于反事实因果推断,区分不同模态的处理变量以减少单模态偏见。
- 在MVSA-Single和MVSA-Multiple数据集上达到新最优性能。
- 适合关注多模态模型公平性与鲁棒性的研究者使用。
随着多媒体技术快速发展,从单模态文本情感分析转向多模态图像-文本情感分析受到学术界和工业界的广泛关注。然而,多模态情感分析易受单模态数据偏见影响,例如文本情感因显式语义误导,导致最终分类准确率偏低。本文提出一种新型反事实多模态情感分析框架(CF-MSA),利用因果反事实推断构建多模态情感因果推理机制。CF-MSA通过区分各模态间的处理变量,缓解单模态偏见的直接效应,并保证模态间异质性。此外,针对模态间信息互补性与偏见差异,设计新的优化目标,有效融合多模态信息并降低各模态固有偏见。在两个公开数据集MVSA-Single和MVSA-Multiple上的实验结果表明,所提CF-MSA具备优异的去偏能力,达到当前最优性能。代码与数据集将开源,以促进后续研究。
原文摘要 · Abstract (English)
With the rapid development of multimedia, the shift from unimodal textual sentiment analysis to multimodal image-text sentiment analysis has obtained academic and industrial attention in recent years. However, multimodal sentiment analysis is affected by unimodal data bias, e.g., text sentiment is misleading due to explicit sentiment semantic, leading to low accuracy in the final sentiment classification. In this paper, we propose a novel CounterFactual Multimodal Sentiment Analysis framework (CF-MSA) using causal counterfactual inference to construct multimodal sentiment causal inference. CF-MSA mitigates the direct effect from unimodal bias and ensures heterogeneity across modalities by differentiating the treatment variables between modalities. In addition, considering the information complementarity and bias differences between modalities, we propose a new optimisation objective to effectively integrate different modalities and reduce the inherent bias from each modality. Experimental results on two public datasets, MVSA-Single and MVSA-Multiple, demonstrate that the proposed CF-MSA has superior debiasing capability and achieves new state-of-the-art performances. We will release the code and datasets to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。