arXiv:2412.07909cs.LGcs.AI2024-12被引 24

揭示对比多模态学习中模态差异的成因并提出缓解方法

Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning

  • 通过梯度流动分析发现数据错配与温度参数是模态差距主因
  • 改进温度调度与模态交换可缩小模态间隙,提升图像文本检索性能
  • 为理解CLIP类模型提供理论支持,适合多模态研究者参考

多模态学习近年来广受欢迎,在零样本分类及感知生成任务中表现优异。以对比语言-图像预训练(CLIP)为代表的模型通过对比学习构建共享表示空间,以融合图像与文本等不同模态。然而,其内在工作机制仍不清晰。值得注意的是,这些模型常出现模态间隙,即不同模态在共享空间中分布分离。本文深入分析模态间隙的产生机制,通过刻画梯度流学习动态,识别出数据错配对与可学习温度参数在训练过程中导致并维持模态间隙的关键作用。进一步地,理论洞察在实际CLIP模型上得到实验验证。基于此,我们提出包括合理温度调度和模态交换在内的缓解策略。结果表明,缩小模态间隙能有效提升图像-文本检索等任务的性能。

原文摘要 · Abstract (English)

Multimodal learning has recently gained significant popularity, demonstrating impressive performance across various zero-shot classification tasks and a range of perceptive and generative applications. Models such as Contrastive Language-Image Pretraining (CLIP) are designed to bridge different modalities, such as images and text, by learning a shared representation space through contrastive learning. Despite their success, the working mechanisms underlying multimodal learning are not yet well understood. Notably, these models often exhibit a modality gap, where different modalities occupy distinct regions within the shared representation space. In this work, we conduct an in-depth analysis of the emergence of modality gap by characterizing the gradient flow learning dynamics. Specifically, we identify the critical roles of mismatched data pairs and a learnable temperature parameter in causing and perpetuating the modality gap during training. Furthermore, our theoretical insights are validated through experiments on practical CLIP models. These findings provide principled guidance for mitigating the modality gap, including strategies such as appropriate temperature scheduling and modality swapping. Additionally, we demonstrate that closing the modality gap leads to improved performance on tasks such as image-text retrieval.

多模态学习对比学习模态间隙CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。