arXiv:2502.15682cs.CV2025-02中稿 · CBMI 2025被引 10

用文本生成视觉提示,提升图像检索准确率。

ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval

  • 用MLP将文本转为视觉提示,动态优化图像编码。
  • 在CLIP/SigLIP上提升检索性能,零样本泛化更强。
  • 适合需要快速适配新场景的图像检索应用。

本文旨在提升文本到图像检索的性能。为此,提出一种新框架ELIP(增强型语言-图像预训练),通过简单MLP将文本查询映射为一组视觉提示,用于条件化ViT图像编码,从而增强大规模预训练视觉-语言模型的检索能力。ELIP可轻松应用于CLIP、SigLIP和BLIP-2等主流模型。为在有限算力下训练,采用全局难例挖掘与大规模数据集构建策略。评估方面,建立两个新的分布外(OOD)基准:遮挡版COCO和ImageNet-R,以测试模型在不同领域下的零样本泛化能力。实验表明,ELIP显著提升CLIP/SigLIP/SigLIP-2的文本到图像检索性能,在多个基准上超越BLIP-2,并提供简便方法适应分布外数据集。

原文摘要 · Abstract (English)

The objective in this paper is to improve the performance of text-to-image retrieval. To this end, we introduce a new framework that can boost the performance of large-scale pre-trained vision-language models, so that they can be used for text-to-image re-ranking. The approach, Enhanced Language-Image Pre-training (ELIP), uses the text query, via a simple MLP mapping network, to predict a set of visual prompts to condition the ViT image encoding. ELIP can easily be applied to the commonly used CLIP, SigLIP and BLIP-2 networks. To train the architecture with limited computing resources, we develop a 'student friendly' best practice, involving global hard sample mining, and curation of a large-scale dataset. On the evaluation side, we set up two new out-of-distribution (OOD) benchmarks, Occluded COCO and ImageNet-R, to assess the zero-shot generalisation of the models to different domains. The results demonstrate that ELIP significantly boosts CLIP/SigLIP/SigLIP-2 text-to-image retrieval performance and outperforms BLIP-2 on several benchmarks, as well as providing an easy means to adapt to OOD datasets.

图像检索视觉语言预训练模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。