arXiv:2508.04028cs.CVcs.IR2025-08被引 1

通过双提示动态优化语义与视觉特征,提升图像文本检索精度。

Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval

  • 设计双提示框架,分别优化属性与类别描述权重。
  • 在超过1500个细粒度类别上实现领先检索性能。
  • 适合需要精准图文匹配的下游应用研究者。

提示学习在适配预训练视觉-语言模型(VLMs)至下游任务中表现卓越,但在图像-文本检索(ITR)任务中仍具挑战性。我们发现,难点在于区分细粒度属性与相似子类别。为此,提出双提示学习与联合类别-属性重加权框架(DCAR),通过动态调整语义与视觉维度的提示向量,提升CLIP在下游ITR任务中的表现。该方法在属性层面根据文本-图像互信息相关性动态更新属性描述权重;在类别层面引入多视角负样本并采用类别匹配加权策略以学习子类别差异。为验证方法,构建了细粒度描述检索数据集(FDRD),涵盖超过1500个细粒度类别与23万张带属性标注的图文对。大量实验表明,DCAR在FDRD上优于现有基线,达到当前最优性能。

原文摘要 · Abstract (English)

Recently, prompt learning has demonstrated remarkable success in adapting pre-trained Vision-Language Models (VLMs) to various downstream tasks such as image classification. However, its application to the downstream Image-Text Retrieval (ITR) task is more challenging. We find that the challenge lies in discriminating both fine-grained attributes and similar subcategories of the downstream data. To address this challenge, we propose Dual prompt Learning with Joint Category-Attribute Reweighting (DCAR), a novel dual-prompt learning framework to achieve precise image-text matching. The framework dynamically adjusts prompt vectors from both semantic and visual dimensions to improve the performance of CLIP on the downstream ITR task. Based on the prompt paradigm, DCAR jointly optimizes attribute and class features to enhance fine-grained representation learning. Specifically, (1) at the attribute level, it dynamically updates the weights of attribute descriptions based on text-image mutual information correlation; (2) at the category level, it introduces negative samples from multiple perspectives with category-matching weighting to learn subcategory distinctions. To validate our method, we construct the Fine-class Described Retrieval Dataset (FDRD), which serves as a challenging benchmark for ITR in downstream data domains. It covers over 1,500 downstream fine categories and 230,000 image-caption pairs with detailed attribute annotations. Extensive experiments on FDRD demonstrate that DCAR achieves state-of-the-art performance over existing baselines.

视觉语言模型图像检索提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。