arXiv:2503.04839cs.CVcs.AI2025-03中稿 · ICLR被引 6

让视觉语言模型更聪明地选示范样本,提升多模态推理能力

Advancing Multimodal In-Context Learning in Large Vision-Language Models with Task-aware Demonstrations

  • 基于任务感知注意力,自动筛选并排序示范样本
  • 在9个数据集上显著提升5种大模型的多模态推理表现
  • 适合需要精准提示设计的应用场景,如智能客服、医疗分析

多模态上下文学习(ICL)已成为大型视觉语言模型(LVLMs)的关键能力,但其性能受图像-文本输入复杂性和输入配置敏感性影响较大。本文揭示了多模态ICL的核心机制,指出任务映射是构建鲁棒示范序列的关键。基于此,提出轻量级解码器模型SabER,通过任务感知注意力,以自回归方式从示范库中智能选择和排列示范样本,实现细粒度特征提取与跨模态推理,迭代优化任务映射以生成高质量示范序列。在覆盖五种LVLMs和九个基准数据集的实验中,SabER不仅展现出强大实证性能,还深化了对任务语义与多模态示范之间交互的理解。研究强调了有原则的示范序列配置的重要性,并为实际应用中的多模态ICL提升开辟新路径。

原文摘要 · Abstract (English)

Multimodal in-context learning (ICL) has emerged as a key capability of Large Vision-Language Models (LVLMs), driven by their increasing scale and applicability. Despite its promise, effective ICL in the multimodal setting remains challenging due to the inherent complexity of image-text inputs and the high sensitivity of ICL performance to input configurations. In this work, we shed light on the core mechanism underlying multimodal ICL, identifying task mapping as a crucial factor in configuring robust in-context demonstration (ICD) sequences. Building on these insights, we propose \textit{SabER}, a lightweight yet powerful decoder-only transformer equipped with task-aware attention, which intelligently selects and arranges ICDs from a demonstration library in an autoregressive fashion. This design enables fine-grained feature extraction and cross-modal reasoning, iteratively refining task mapping to generate high-quality ICD sequences. Through extensive experiments covering five LVLMs and nine benchmark datasets, SabER not only demonstrates strong empirical performance, but also provides deeper understanding of how task semantics interact with multimodal ICDs. Our findings highlight the importance of principled ICD sequence configuration and open new avenues to enhance multimodal ICL in a wide range of real-world scenarios.

多模态学习提示工程视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。