arXiv:2511.22664cs.CV2025-11NeurIPS被引 5

让视觉语言模型根据输入自动生成个性化提示,提升少样本与跨域任务表现。

VaMP: Variational Multi-Modal Prompt Learning for Vision-Language Models

  • 通过变分推断生成随样本变化的提示,捕捉输入差异和不确定性。
  • 在少样本分类与跨域泛化任务上达到当前最优性能,显著优于固定提示方法。
  • 适合需要适应多变场景的视觉语言模型微调研究者使用。

视觉语言模型(如CLIP)在零样本设置下表现出强大泛化能力,但在有限监督下的下游任务适配仍具挑战。现有多模态提示学习方法通常依赖固定共享提示和确定性参数,难以捕捉实例级差异或跨任务/领域的不确定性。为此,我们提出一种新型变分多模态提示学习(VaMP)框架,实现样本相关的、不确定性感知的多模态表示学习中的提示调优。VaMP通过从学习到的后验分布中采样生成实例条件提示,使模型能根据输入内容个性化响应。为进一步融合局部与全局语义,引入基于实例表征与类别原型的类感知先验。在此基础上,将提示调优建模为对潜在提示表示的变分推断,并通过重参数化采样实现端到端训练。在少样本与领域泛化基准上的实验表明,VaMP达到当前最优性能,验证了建模不确定性与任务结构的优越性。

原文摘要 · Abstract (English)

Vision-language models (VLMs), such as CLIP, have shown strong generalization under zero-shot settings, yet adapting them to downstream tasks with limited supervision remains a significant challenge. Existing multi-modal prompt learning methods typically rely on fixed, shared prompts and deterministic parameters, which limits their ability to capture instance-level variation or model uncertainty across diverse tasks and domains. To tackle this issue, we propose a novel Variational Multi-Modal Prompt Learning (VaMP) framework that enables sample-specific, uncertainty-aware prompt tuning in multi-modal representation learning. VaMP generates instance-conditioned prompts by sampling from a learned posterior distribution, allowing the model to personalize its behavior based on input content. To further enhance the integration of local and global semantics, we introduce a class-aware prior derived from the instance representation and class prototype. Building upon these, we formulate prompt tuning as variational inference over latent prompt representations and train the entire framework end-to-end through reparameterized sampling. Experiments on few-shot and domain generalization benchmarks show that VaMP achieves state-of-the-art performance, highlighting the benefits of modeling both uncertainty and task structure in our method. Project page: https://visual-ai.github.io/vamp

多模态提示变分推断少样本学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。