arXiv:2510.24331cs.LGcs.CV2025-10被引 3

探究视觉语言模型如何利用上下文示例学习,发现其主要依赖文本而非图像。

What do vision-language models see in the context? Investigating multimodal in-context learning

论文配图:What do vision-language models see in the context? Investigating multimodal in-context learning
图 1 · 摘自论文原文
  • 对比7种模型在3个图像描述任务上的表现,分析提示设计与架构影响。
  • 随着示例增多,模型注意力仍集中在文本上,视觉信息未被有效利用。
  • 指令微调虽提升任务理解,但削弱对上下文示例的依赖,存在权衡。

上下文学习(ICL)使大语言模型无需参数更新即可从示例中学习任务。尽管该技术在大语言模型中已得到广泛研究,但在视觉语言模型(VLMs)中的有效性仍缺乏探索。本文系统研究了VLMs的ICL,评估了7种跨越4种架构的模型在3个图像描述基准上的表现。我们分析了提示设计、架构选择和训练策略对多模态ICL的影响。据我们所知,首次揭示了随着上下文示例数量增加,VLM中注意力模式的变化。结果表明,在图文交错数据上训练可提升ICL性能,但并不意味着能有效整合示例中的视觉与文本信息;而指令微调虽改善指令遵循能力,却可能降低对上下文示例的依赖,反映出指令对齐与上下文适应之间的权衡。注意力分析进一步显示,当前VLMs主要关注文本线索,未能充分利用视觉信息,表明其多模态整合能力有限。这些发现揭示了现有VLMs在上下文学习方面的关键局限,并为提升其从多模态示例中学习的能力提供了启示。

原文摘要 · Abstract (English)

In-context learning (ICL) enables Large Language Models (LLMs) to learn tasks from demonstration examples without parameter updates. Although it has been extensively studied in LLMs, its effectiveness in Vision-Language Models (VLMs) remains underexplored. In this work, we present a systematic study of ICL in VLMs, evaluating seven models spanning four architectures on three image captioning benchmarks. We analyze how prompt design, architectural choices, and training strategies influence multimodal ICL. To our knowledge, we are the first to analyze how attention patterns in VLMs vary with an increasing number of in-context demonstrations. Our results reveal that training on imag-text interleaved data enhances ICL performance but does not imply effective integration of visual and textual information from demonstration examples. In contrast, instruction tuning improves instruction-following but can reduce reliance on in-context demonstrations, suggesting a trade-off between instruction alignment and in-context adaptation. Attention analyses further show that current VLMs primarily focus on textual cues and fail to leverage visual information, suggesting a limited capacity for multimodal integration. These findings highlight key limitations in the ICL abilities of current VLMs and provide insights for enhancing their ability to learn from multimodal in-context examples.

视觉语言模型上下文学习注意力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。