arXiv:2510.03268cs.LGcs.AI2025-10被引 10

揭示多模态对比学习中模态差距的成因与影响

Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment

  • 理论证明模态间隙源于维度坍缩,约束下可收敛至零或最小夹角
  • 在子空间约束下样本对无法完全对齐,影响下游任务性能
  • 提出通过超平面旋转和共享空间投影实现完美对齐,适用于多模态模型优化

多模态对比学习(MCL)旨在将不同模态的数据嵌入共享的嵌入空间。然而,实证发现不同模态的表示占据嵌入空间中完全分离的区域,这一现象称为模态间隙。且现有实验对模态间隙大小与下游性能的关系结论不一。本文首次建立分析MCL收敛最优表示与模态对齐的理论框架。证明:无约束或在锥约束下,模态间隙收敛至零;在子空间约束(即模态表示落入两个不同超平面,由维度坍缩导致)下,模态间隙收敛至两超平面间的最小夹角。该结果表明,维度坍缩是模态间隙的根本原因。此外,定理显示在子空间约束下,样本对无法完全对齐。模态间隙通过影响样本对间对齐程度来影响下游性能。我们进一步证明,可通过超平面旋转或共享空间投影实现两模态间的完美对齐。

原文摘要 · Abstract (English)

Multimodal contrastive learning (MCL) aims to embed data from different modalities in a shared embedding space. However, empirical evidence shows that representations from different modalities occupy completely separate regions of embedding space, a phenomenon referred to as the modality gap. Moreover, experimental findings on how the size of the modality gap influences downstream performance are inconsistent. These observations raise two key questions: (1) What causes the modality gap? (2) How does it affect downstream tasks? To address these questions, this paper introduces the first theoretical framework for analyzing the convergent optimal representations of MCL and the modality alignment when training is optimized. Specifically, we prove that without any constraint or under the cone constraint, the modality gap converges to zero. Under the subspace constraint (i.e., representations of two modalities fall into two distinct hyperplanes due to dimension collapse), the modality gap converges to the smallest angle between the two hyperplanes. This result identifies \emph{dimension collapse} as the fundamental origin of the modality gap. Furthermore, our theorems demonstrate that paired samples cannot be perfectly aligned under the subspace constraint. The modality gap influences downstream performance by affecting the alignment between sample pairs. We prove that, in this case, perfect alignment between two modalities can still be achieved via two ways: hyperplane rotation and shared space projection.

多模态对比学习模态对齐理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。