arXiv:2509.21805cs.CL2025-09被引 1

提出新模型提升多模态理解的因果推理能力,避免数据偏差影响

Towards Minimal Causal Representations for Human Multimodal Language Understanding

  • 用信息瓶颈过滤无关噪声,分离因果与捷径特征
  • 在情感、幽默、反讽任务中显著提升跨分布泛化性能
  • 适合关注可解释性与鲁棒性的多模态研究者

人类多模态语言理解(MLU)旨在通过融合异构模态线索推断人类意图。现有方法多采用“学习注意力”范式,最大化数据与标签间的互信息以提升预测性能,但易受数据集偏差影响,导致模型混淆统计捷径与真实因果特征,损害分布外(OOD)泛化能力。为此,我们提出因果多模态信息瓶颈(CaMIB)模型,基于因果原则而非传统似然。首先应用信息瓶颈过滤单模态输入中的无关噪声;参数化掩码生成器将融合的多模态表示解耦为因果与捷径子表示;为保证因果特征全局一致性,引入工具变量约束,并通过随机重组因果与捷径特征实施后门调整以稳定因果估计。在多模态情感分析、幽默检测和反讽检测任务及多个分布外测试集上的实验表明,CaMIB效果显著。理论与实证分析进一步验证其可解释性与合理性。

原文摘要 · Abstract (English)

Human Multimodal Language Understanding (MLU) aims to infer human intentions by integrating related cues from heterogeneous modalities. Existing works predominantly follow a ``learning to attend" paradigm, which maximizes mutual information between data and labels to enhance predictive performance. However, such methods are vulnerable to unintended dataset biases, causing models to conflate statistical shortcuts with genuine causal features and resulting in degraded out-of-distribution (OOD) generalization. To alleviate this issue, we introduce a Causal Multimodal Information Bottleneck (CaMIB) model that leverages causal principles rather than traditional likelihood. Concretely, we first applies the information bottleneck to filter unimodal inputs, removing task-irrelevant noise. A parameterized mask generator then disentangles the fused multimodal representation into causal and shortcut subrepresentations. To ensure global consistency of causal features, we incorporate an instrumental variable constraint, and further adopt backdoor adjustment by randomly recombining causal and shortcut features to stabilize causal estimation. Extensive experiments on multimodal sentiment analysis, humor detection, and sarcasm detection, along with OOD test sets, demonstrate the effectiveness of CaMIB. Theoretical and empirical analyses further highlight its interpretability and soundness.

多模态因果推理可解释性泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。