提出多模态提示调优,让大模型零样本学会跨模态任务
M$^2$PT: Multimodal Prompt Tuning for Zero-shot Instruction Learning
- 在视觉编码器和语言处理器中分别注入图文提示
- 在多个数据集上超越现有方法,零样本性能更优
- 适合需要高效微调多模态模型的研究者
多模态大语言模型在多个领域表现优异,提升其对未见任务的零样本泛化能力成为关键。指令微调是实现零样本泛化的有效策略,通过在多样化多模态任务上微调预训练模型。随着多模态大模型规模持续扩大,参数高效微调愈发重要。然而,现有大多数高效方法仅关注单模态,忽略微调过程中的多模态特性。本文提出一种新型多模态提示调优(M$^2$PT)方法,用于多模态大模型的高效指令微调。M$^2$PT在微调过程中将视觉提示和文本提示分别融入视觉编码器与语言处理器,促进跨模态特征的提取与对齐。在多个多模态评估数据集上的实验表明,该方法性能优于多种先进基线。全面的消融实验证实了提示设计的有效性及方法的高效性。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) demonstrate remarkable performance across a wide range of domains, with increasing emphasis on enhancing their zero-shot generalization capabilities for unseen tasks across various modalities. Instruction tuning has emerged as an effective strategy for achieving zero-shot generalization by finetuning pretrained models on diverse multimodal tasks. As the scale of MLLMs continues to grow, parameter-efficient finetuning becomes increasingly critical. However, most existing parameter-efficient approaches focus only on single modalities and often overlook the multimodal characteristics during finetuning. In this work, we introduce a novel Multimodal Prompt Tuning (M$^2$PT) approach for efficient instruction tuning of MLLMs. M$^2$PT effectively integrates visual and textual prompts into the vision encoder and language processor respectively during finetuning, facilitating the extraction and alignment of features across modalities. Empirical results on various multimodal evaluation datasets demonstrate the superior performance of our approach compared to several state-of-the-art baselines. A comprehensive set of ablation studies validates the effectiveness of our prompt design and the efficiency of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。