arXiv:2505.03703cs.CVcs.LG2025-05被引 9

量化并缩小图文表示学习中的模态差距,提升跨模态任务性能

Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning

  • 提出基于谱分析和最优传输的新度量方法
  • 实验证明可显著减小图文嵌入的分离程度
  • 适合关注跨模态对齐与下游应用的研究者

视觉-语言模型(VLMs)可将文本和图像映射到共享表示空间,但研究发现这些模型存在模态差距现象——不同模态的嵌入在空间中明显分离。这种错位会损害多模态检索、聚类及零样本分类等下游任务表现。然而,此前缺乏通用且实用的方法来精确评估并缓解该问题。本文提出新的度量方法(基于谱分析与最优传输)及有效技术,通过在多个图像-文本数据集(如 COCO、Flickr30k)和模型上进行广泛实验,验证了其有效性,并显著提升了下游任务性能。代码已公开。

原文摘要 · Abstract (English)

Vision-language models (VLMs) allow to embed texts and images in a shared representation space. However, it has been shown that these models are subject to a modality gap phenomenon meaning there exists a clear separation between the embeddings from one modality and another in the embedding space. While this misalignment is detrimental for downstream tasks such as multimodal retrieval, multimodal clustering or zero-shot classification, etc. no generic and practical methods have so far been proposed to assess it precisely and even reduce it. We therefore propose novel measures and effective techniques (spectral- and optimal transport-based methods) to achieve this goal. Extensive experiments conducted on several image-text datasets and models demonstrate their effectiveness and beneficial effects on downstream tasks. Our code is available at the URL provided in the paper's abstract.

模态对齐图文表示嵌入空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。