arXiv:2501.07769cs.LGcs.CV2025-01IJCAI被引 2

提出双向模态交互提示学习,提升视觉语言模型对齐能力

BMIP: Bi-directional Modality Interaction Prompt Learning for VLM

  • 通过动态加权视觉与语言模态信息增强跨模态对齐
  • 在三种评估范式下均超越现有最佳方法
  • 可与其它提示方法结合,适合需要多模态适配的场景

视觉语言模型(VLM)展现出强大的泛化能力,提示学习因其能将预训练的VLM适配到特定下游任务而受到广泛关注。然而,现有研究主要聚焦于单模态提示或单向模态交互,忽视了视觉与语言模态间交互带来的强大对齐效应。为此,我们提出一种新型提示学习方法——双向模态交互提示(BMIP),通过学习注意力层的信息,动态加权双模态信息,相比简单的信息聚合方法,提升了可训练性和模态间一致性。为评估提示学习方法的有效性,我们提出一种更贴近现实的评估范式——开放世界泛化,以补充广泛采用的跨数据集迁移和领域泛化任务。在多个数据集上的综合实验表明,BMIP不仅在所有三种评估范式中均优于当前最先进方法,且具备良好灵活性,可与其他基于提示的方法结合,实现一致的性能提升。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have exhibited remarkable generalization capabilities, and prompt learning for VLMs has attracted great attention for the ability to adapt pre-trained VLMs to specific downstream tasks. However, existing studies mainly focus on single-modal prompts or uni-directional modality interaction, overlooking the powerful alignment effects resulting from the interaction between the vision and language modalities. To this end, we propose a novel prompt learning method called $\underline{\textbf{B}}i-directional \underline{\textbf{M}}odality \underline{\textbf{I}}nteraction \underline{\textbf{P}}rompt (BMIP)$, which dynamically weights bi-modal information through learning the information of the attention layer, enhancing trainability and inter-modal consistency compared to simple information aggregation methods. To evaluate the effectiveness of prompt learning methods, we propose a more realistic evaluation paradigm called open-world generalization complementing the widely adopted cross-dataset transfer and domain generalization tasks. Comprehensive experiments on various datasets reveal that BMIP not only outperforms current state-of-the-art methods across all three evaluation paradigms but is also flexible enough to be combined with other prompt-based methods for consistent performance enhancement.

视觉语言模型提示学习跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。