提出能量模型生成多提示,提升视觉语言模型泛化能力
A Retrospect to Multi-prompt Learning across Vision and Language
- 用能量模型隐式生成多个提示嵌入
- 实证与理论证明多提示优于单提示
- 兼顾领域内与开放词汇泛化,参数高效
视觉语言预训练模型(VLMs)推动了视觉领域前所未有的进展。提示学习作为访问VLMs的利器,可实现下游任务的快速适应。尽管现有研究集中于单提示范式,对多提示学习的技术潜力关注甚少。本文系统回顾了跨模态多提示学习,将近期发现的恒定模态差距现象扩展至可学习提示,并从实证与理论上验证了多提示增强在视觉语言迁移中的优越性。基于此,我们提出能量基础多提示学习(EMPL),通过从由VLMs隐式定义的能量分布中采样实例,生成多个提示嵌入。EMPL不仅参数高效,且严格平衡了领域内与领域外的开放词汇泛化能力。大量实验验证了我们的主张及EMPL的卓越性能。
原文摘要 · Abstract (English)
The vision community is undergoing the unprecedented progress with the emergence of Vision-Language Pretraining Models (VLMs). Prompt learning plays as the holy grail of accessing VLMs since it enables their fast adaptation to downstream tasks with limited resources. Whereas existing researches milling around single-prompt paradigms, rarely investigate the technical potential behind their multi-prompt learning counterparts. This paper aims to provide a principled retrospect for vision-language multi-prompt learning. We extend the recent constant modality gap phenomenon to learnable prompts and then, justify the superiority of vision-language transfer with multi-prompt augmentation, empirically and theoretically. In terms of this observation, we propose an Energy-based Multi-prompt Learning (EMPL) to generate multiple prompt embeddings by drawing instances from an energy-based distribution, which is implicitly defined by VLMs. So our EMPL is not only parameter-efficient but also rigorously lead to the balance between in-domain and out-of-domain open-vocabulary generalization. Comprehensive experiments have been conducted to justify our claims and the excellence of EMPL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。