通过因果注意力分离模态内与跨模态偏见,提升情感分析泛化能力
Disentangling Bias by Modeling Intra- and Inter-modal Causal Attention for Multimodal Sentiment Analysis
- 构建多关系图模型捕捉模态间与模态内依赖
- 分离因果特征与统计捷径特征,显著降低偏差影响
- 适合关注模型鲁棒性与可解释性的多模态研究者
多模态情感分析(MSA)旨在融合文本、音频、视觉等多源信息理解人类情绪。然而,现有方法常受模态内与跨模态虚假相关性干扰,导致模型依赖统计捷径而非真实因果关系,削弱泛化性能。为此,我们提出多关系多模态因果干预(MMCI)框架,利用因果理论中的后门调整机制缓解此类捷径的混杂效应。具体地,首先将多模态输入建模为多关系图,显式捕捉模态内与跨模态依赖;随后通过注意力机制分别估计并解耦对应于这些关系的因果特征与捷径特征;最后应用后门调整,对捷径特征进行分层处理,并动态融合因果特征,以在分布偏移下生成稳定预测。在多个标准MSA数据集及分布外(OOD)测试集上的大量实验表明,该方法有效抑制偏见并提升性能。
原文摘要 · Abstract (English)
Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating information from multiple modalities, such as text, audio, and visual data. However, existing methods often suffer from spurious correlations both within and across modalities, leading models to rely on statistical shortcuts rather than true causal relationships, thereby undermining generalization. To mitigate this issue, we propose a Multi-relational Multimodal Causal Intervention (MMCI) framework, which leverages the backdoor adjustment from causal theory to address the confounding effects of such shortcuts. Specifically, we first model the multimodal inputs as a multi-relational graph to explicitly capture intra- and inter-modal dependencies. Then, we apply an attention mechanism to separately estimate and disentangle the causal features and shortcut features corresponding to these intra- and inter-modal relations. Finally, by applying the backdoor adjustment, we stratify the shortcut features and dynamically combine them with the causal features to encourage MMCI to produce stable predictions under distribution shifts. Extensive experiments on several standard MSA datasets and out-of-distribution (OOD) test sets demonstrate that our method effectively suppresses biases and improves performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。