arXiv:2412.10680cs.CVcs.IR2024-12中稿 · WACV 2025被引 6

动态生成提示词,让模型跨领域检索更灵活高效。

UCDR-Adapter: Exploring Adaptation of Pre-Trained Vision-Language Models for Universal Cross-Domain Retrieval

  • 用可学习的适配器和动态提示生成机制增强视觉语言模型。
  • 在多个跨域检索任务上超越现有方法,最高提升6.3%。
  • 推理时无需文本输入,适合实际应用中的快速检索场景。

通用跨域检索(UCDR)旨在无语义标签条件下从未见领域和类别中检索相关图像,确保强泛化能力。现有方法多采用预训练视觉语言模型进行提示调优,但受限于静态提示,适应性不足。本文提出UCDR-Adapter,通过两阶段训练策略增强预训练模型:第一阶段,源适配器学习利用可学习文本语义模板,结合领域特定视觉知识,通过动量更新与双重损失函数优化类别与领域提示,实现鲁棒对齐;第二阶段,目标提示生成通过关注掩码后的源提示动态构建新提示,实现对未见领域和类别的无缝适应。相比以往方法,UCDR-Adapter能动态适应数据分布变化,提升灵活性与泛化性。推理时仅使用图像分支与生成提示,无需文本输入,实现高效检索。大量基准实验表明,UCDR-Adapter在UCDR、U(c)CDR和U(d)CDR设置下持续优于ProS及其他前沿方法。

原文摘要 · Abstract (English)

Universal Cross-Domain Retrieval (UCDR) retrieves relevant images from unseen domains and classes without semantic labels, ensuring robust generalization. Existing methods commonly employ prompt tuning with pre-trained vision-language models but are inherently limited by static prompts, reducing adaptability. We propose UCDR-Adapter, which enhances pre-trained models with adapters and dynamic prompt generation through a two-phase training strategy. First, Source Adapter Learning integrates class semantics with domain-specific visual knowledge using a Learnable Textual Semantic Template and optimizes Class and Domain Prompts via momentum updates and dual loss functions for robust alignment. Second, Target Prompt Generation creates dynamic prompts by attending to masked source prompts, enabling seamless adaptation to unseen domains and classes. Unlike prior approaches, UCDR-Adapter dynamically adapts to evolving data distributions, enhancing both flexibility and generalization. During inference, only the image branch and generated prompts are used, eliminating reliance on textual inputs for highly efficient retrieval. Extensive benchmark experiments show that UCDR-Adapter consistently outperforms ProS in most cases and other state-of-the-art methods on UCDR, U(c)CDR, and U(d)CDR settings.

跨域检索视觉语言模型动态提示适配器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。