arXiv:2507.07415cs.CV2025-07

提出高效图文分类提示交互方法,显著降低计算开销。

EPIC: Efficient Prompt Interaction for Text-Image Classification

  • 在中间层使用时序提示,通过相似性交互融合多模态信息。
  • 参数量仅需基础模型1%,计算资源消耗大幅减少。
  • 在多个数据集上表现优异,适合资源受限场景应用。

近年来,大规模预训练多模态模型(LMMs)广泛涌现,成功整合视觉与语言模态,在图文分类等任务中取得显著成果。然而,模型规模的增大导致下游任务微调时计算成本过高。为此,本文提出一种新型高效提示交互策略——高效图文分类提示交互(EPIC)。具体而言,在中间层引入时序提示,并通过基于相似性的提示交互机制,促进模态间充分的信息交换。该方法显著降低计算资源消耗与可训练参数量(约为基础模型的1%),同时在UPMC-Food101和SNLI-VE数据集上表现更优,在MM-IMDB上保持相当水平。

原文摘要 · Abstract (English)

In recent years, large-scale pre-trained multimodal models (LMMs) generally emerge to integrate the vision and language modalities, achieving considerable success in multimodal tasks, such as text-image classification. The growing size of LMMs, however, results in a significant computational cost for fine-tuning these models for downstream tasks. Hence, prompt-based interaction strategy is studied to align modalities more efficiently. In this context, we propose a novel efficient prompt-based multimodal interaction strategy, namely Efficient Prompt Interaction for text-image Classification (EPIC). Specifically, we utilize temporal prompts on intermediate layers, and integrate different modalities with similarity-based prompt interaction, to leverage sufficient information exchange between modalities. Utilizing this approach, our method achieves reduced computational resource consumption and fewer trainable parameters (about 1\% of the foundation model) compared to other fine-tuning strategies. Furthermore, it demonstrates superior performance on the UPMC-Food101 and SNLI-VE datasets, while achieving comparable performance on the MM-IMDB dataset.

多模态提示学习高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。