arXiv:2409.04693cs.AI2024-09被引 11

解决视觉语言模型在模态缺失时的表现问题

MuAP: Multi-step Adaptive Prompt Learning for Vision-Language Model with Missing Modality

  • 分步自适应生成多模态提示,逐步对齐不同模态
  • 在多个基准数据集上显著优于现有方法
  • 适合处理实际场景中不完整模态信息的任务

最近,提示学习在多种视觉-语言(VL)任务中表现出色。然而,现有提示模型主要关注完整模态下的提示生成与策略,未能反映现实中部分模态缺失的真实情况。本文首次系统研究了模态不完整时的提示学习行为,发现提示模型对缺失模态极为敏感。为此,我们提出一种新的多步自适应提示学习(MuAP)框架,旨在生成多模态提示并进行多步提示调优,通过迭代对齐模态实现自适应知识学习。具体而言,为每种模态生成提示,并设计策略将其整合到Transformer模型中;随后依次执行单阶段与对齐阶段的提示调优,使每个模态提示能自主、自适应地学习,从而缓解以往仅文本提示可学习导致的不平衡问题。大量实验表明,该方法在所有基准数据集上均显著优于当前最先进水平。

原文摘要 · Abstract (English)

Recently, prompt learning has garnered considerable attention for its success in various Vision-Language (VL) tasks. However, existing prompt-based models are primarily focused on studying prompt generation and prompt strategies with complete modality settings, which does not accurately reflect real-world scenarios where partial modality information may be missing. In this paper, we present the first comprehensive investigation into prompt learning behavior when modalities are incomplete, revealing the high sensitivity of prompt-based models to missing modalities. To this end, we propose a novel Multi-step Adaptive Prompt Learning (MuAP) framework, aiming to generate multimodal prompts and perform multi-step prompt tuning, which adaptively learns knowledge by iteratively aligning modalities. Specifically, we generate multimodal prompts for each modality and devise prompt strategies to integrate them into the Transformer model. Subsequently, we sequentially perform prompt tuning from single-stage and alignment-stage, allowing each modality-prompt to be autonomously and adaptively learned, thereby mitigating the imbalance issue caused by only textual prompts that are learnable in previous works. Extensive experiments demonstrate the effectiveness of our MuAP and this model achieves significant improvements compared to the state-of-the-art on all benchmark datasets

提示学习多模态自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。