arXiv:2604.18460cs.LG2026-04ACL

从因果视角学习跨环境稳定的多模态表示,提升模型鲁棒性。

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective

论文配图:Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective
图 1 · 摘自论文原文
  • 基于因果推断分离出因果不变表示与环境特异噪声
  • 在多个数据集上实现当前最优性能,尤其在分布外和噪声数据上表现突出
  • 适合需要高泛化能力的多模态情感计算任务

多模态情感计算旨在通过语言、语音和视觉模态预测人类的情感、情绪、意图和观点。然而,现有模型常学习到虚假相关性,在分布偏移或噪声模态下损害泛化能力。为此,我们提出因果模态不变表示(CmIR)学习框架,从因果推断角度出发,引入理论支撑的解耦方法,将每种模态分解为‘因果不变表示’和‘环境特异的虚假表示’。CmIR通过不变性约束、互信息约束和重建约束,确保所学不变表示在不同环境下保持与标签的稳定预测关系,同时保留原始输入的充分信息。在多个多模态基准上的实验表明,CmIR达到当前最优性能,尤其在分布外数据和噪声数据上表现优异,验证了其鲁棒性和泛化能力。

原文摘要 · Abstract (English)

Multimodal affective computing aims to predict humans' sentiment, emotion, intention, and opinion using language, acoustic, and visual modalities. However, current models often learn spurious correlations that harm generalization under distribution shifts or noisy modalities. To address this, we propose a causal modality-invariant representation (CmIR) learning framework for robust multimodal learning. At its core, we introduce a theoretically grounded disentanglement method that separates each modality into `causal invariant representation' and `environment-specific spurious representation' from a causal inference perspective. CmIR ensures that the learned invariant representations retain stable predictive relationships with labels across different environments while preserving sufficient information from the raw inputs via invariance constraint, mutual information constraint, and reconstruction constraint. Experiments across multiple multimodal benchmarks demonstrate that CmIR achieves state-of-the-art performance. CmIR particularly excels on out-of-distribution data and noisy data, confirming its robustness and generalizability.

多模态因果推理表示学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。