arXiv:2507.08410cs.CV2025-07被引 1

通过双向引导生成视觉语义提示,提升模型对新类别泛化能力

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models

  • 用多模态大模型动态生成带细粒度语义的条件提示
  • 在14个数据集上超越现有最优方法,显著提升新类别识别效果
  • 适合需要强泛化能力的视觉语言模型下游任务

提示学习可高效适配视觉语言模型(VLM)至各类下游任务,但面临两大挑战:一是对未见类别嵌入分布建模不足,导致新类别泛化性能不佳;二是现有方法多将跨模态对齐局限于编码器输出层,难以保持与预训练多模态嵌入空间的拓扑一致性。为此,我们提出MuGCP(多模态互引导条件提示学习),一种用于条件提示生成的新范式。MuGCP利用多模态大语言模型(MLLM)作为条件提示学习器,自适应生成包含丰富高层语义知识的语义条件提示(SCP),用于图像实例。为确保视觉与语义信息在多模态空间中的有效对齐与交互,引入注意力互引导(AMG)模块,通过双向引导生成视觉条件提示(VCP),提升多模态任务性能。此外,提出多提示融合(MPF)机制,将SCP、VCP与上下文提示融合,实现不同提示间无缝协调,增强类别嵌入与实例特异性知识建模。实验表明,MuGCP在14个不同数据集上优于现有最先进方法。

原文摘要 · Abstract (English)

Prompt learning facilitates the efficient adaptation of Vision-Language Models (VLMs) to various downstream tasks. However, it faces two significant challenges: (1) inadequate modeling of class embedding distributions for unseen instances, leading to suboptimal generalization on novel classes; (2) prevailing methodologies predominantly confine cross-modal alignment to the final output layer of vision and text encoders, which fundamentally limits their capacity to preserve topological consistency with pre-trained multi-modal embedding spaces. To this end, we introduce MuGCP (Multi-modal Mutual-Guidance Conditional Prompt Learning), a novel paradigm designed for conditional prompt generation. MuGCP leverages Multi-modal Large Language Models (MLLMs) as conditional prompt learners to adaptively generate Semantic Conditional Prompts (SCP) that incorporate rich, fine-grained high-level semantic knowledge for image instances. To ensure effective alignment and interaction across the multi-modal space of Vision-Language Models (VLMs), we introduce the Attention Mutual-Guidance (AMG) module, which facilitates interactions between visual and semantic information. Through mutual guidance, the AMG module generates Visual Conditional Prompts (VCP), enhancing the model's performance in multi-modal tasks. Additionally, we present a Multi-Prompt Fusion (MPF) mechanism that integrates SCP and VCP with contextual prompts, ensuring seamless coordination among the different prompts and enhancing the modeling of class embeddings and instance-specific knowledge. Our MuGCP outperforms existing state-of-the-art methods on 14 different datasets. The code will be made available after publication.

提示学习多模态视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。