arXiv:2504.10143cs.LGcs.CV2025-04NeurIPS被引 11

揭示跨模态错位在多模态学习中的双重作用,指导实际系统设计。

On the Value of Cross-Modal Misalignment in Multimodal Representation Learning

  • 用潜变量模型定义选择偏差与扰动偏差导致的错位机制。
  • 理论证明对比学习只保留对两类偏差不变的语义信息。
  • 为真实场景下多模态模型设计提供可操作的实践指南。

多模态表示学习(如基于图像-文本对的多模态对比学习,MMCL)旨在通过跨模态对齐学习强表示。该方法依赖于一个核心假设:图像-文本样本是同一概念的两种表征。然而,现实数据集常存在跨模态错位。学界对此有两种观点:一为缓解错位,二为利用错位。本文试图调和二者,为实践者提供指导。通过潜变量模型,我们形式化了两类错位机制:选择偏差(部分语义变量缺失于文本)与扰动偏差(语义变量被改变),二者均导致数据对错位。理论分析表明,在温和假设下,MMCL学习到的表示仅捕获对选择与扰动偏差不变的语义变量信息。这提供了统一理解错位的视角。进一步,我们提出如何将错位影响融入真实机器学习系统设计。通过合成数据与真实图像-文本数据集的广泛实证研究,验证了理论发现,揭示了跨模态错位对多模态表示学习的微妙影响。

原文摘要 · Abstract (English)

Multimodal representation learning, exemplified by multimodal contrastive learning (MMCL) using image-text pairs, aims to learn powerful representations by aligning cues across modalities. This approach relies on the core assumption that the exemplar image-text pairs constitute two representations of an identical concept. However, recent research has revealed that real-world datasets often exhibit cross-modal misalignment. There are two distinct viewpoints on how to address this issue: one suggests mitigating the misalignment, and the other leveraging it. We seek here to reconcile these seemingly opposing perspectives, and to provide a practical guide for practitioners. Using latent variable models we thus formalize cross-modal misalignment by introducing two specific mechanisms: Selection bias, where some semantic variables are absent in the text, and perturbation bias, where semantic variables are altered -- both leading to misalignment in data pairs. Our theoretical analysis demonstrates that, under mild assumptions, the representations learned by MMCL capture exactly the information related to the subset of the semantic variables invariant to selection and perturbation biases. This provides a unified perspective for understanding misalignment. Based on this, we further offer actionable insights into how misalignment should inform the design of real-world ML systems. We validate our theoretical findings via extensive empirical studies on both synthetic data and real image-text datasets, shedding light on the nuanced impact of cross-modal misalignment on multimodal representation learning.

多模态表示学习错位分析理论建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。