arXiv:2409.12408cs.CLcs.MM2024-09被引 1

通过最小化互信息,消除多模态序列中的冗余信息,提升模型泛化能力。

Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences

  • 设计联合学习单一跨模态无关表示的解耦框架
  • 利用互信息最小化约束消除非线性相关性,减少表示冗余
  • 借助无标签数据缓解标注不足问题,增强模型泛化性能

未对齐多模态语言序列的核心挑战在于如何有效融合不同模态信息以获得精细的多模态联合表示。近期提出的解耦与融合方法通过显式学习跨模态无关和模态特定表示,并将其融合为多模态联合表示,取得了良好效果。然而,这些方法通常独立为各模态学习跨模态无关表示,并使用正交约束降低其与模态特定表示之间的线性相关性,忽略了非线性相关性的消除。导致最终的多模态联合表示存在信息冗余,引发过拟合并降低模型泛化能力。本文提出一种基于互信息的表示解耦方法(MIRD),设计新颖的解耦框架,联合学习单一跨模态无关表示。同时引入互信息最小化约束,确保表示解耦效果,从而消除多模态联合表示中的信息冗余。此外,通过引入无标签数据缓解因标注数据有限导致的互信息估计难题,同时帮助揭示多模态数据潜在结构,进一步防止过拟合并提升模型性能。在多个主流基准数据集上的实验结果验证了所提方法的有效性。

原文摘要 · Abstract (English)

The key challenge in unaligned multimodal language sequences lies in effectively integrating information from various modalities to obtain a refined multimodal joint representation. Recently, the disentangle and fuse methods have achieved the promising performance by explicitly learning modality-agnostic and modality-specific representations and then fusing them into a multimodal joint representation. However, these methods often independently learn modality-agnostic representations for each modality and utilize orthogonal constraints to reduce linear correlations between modality-agnostic and modality-specific representations, neglecting to eliminate their nonlinear correlations. As a result, the obtained multimodal joint representation usually suffers from information redundancy, leading to overfitting and poor generalization of the models. In this paper, we propose a Mutual Information-based Representations Disentanglement (MIRD) method for unaligned multimodal language sequences, in which a novel disentanglement framework is designed to jointly learn a single modality-agnostic representation. In addition, the mutual information minimization constraint is employed to ensure superior disentanglement of representations, thereby eliminating information redundancy within the multimodal joint representation. Furthermore, the challenge of estimating mutual information caused by the limited labeled data is mitigated by introducing unlabeled data. Meanwhile, the unlabeled data also help to characterize the underlying structure of multimodal data, consequently further preventing overfitting and enhancing the performance of the models. Experimental results on several widely used benchmark datasets validate the effectiveness of our proposed approach.

多模态表示解耦互信息序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。