arXiv:2506.07936cs.CVcs.AI2025-06被引 6

研究发现视觉语言模型在多模态上下文学习中更依赖复制答案而非真正理解任务。

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

  • 为提升理解力,给每个示例添加生成的推理过程来增强多模态上下文学习
  • 增加演示数量反而降低性能,模型更倾向直接复制而非学习
  • 适用于评估视觉语言模型是否真具备推理能力的研究者

视觉语言模型(VLMs)常被认为具备上下文学习(ICL)能力,类似纯语言模型。尽管有研究指出它们可实现多模态上下文学习(MM-ICL),但多数表现依赖浅层启发式策略,如复制或多数投票,而非真正理解任务。本文通过在分布偏移条件下评估模型——支持示例来自与查询不同的数据集——发现,随着演示数量增加,性能往往下降,模型倾向于复制答案而非从中学习。为此,我们提出一种新的带推理的多模态上下文学习(MM-ICL with Reasoning)流程,为每个演示附加生成的推理过程。我们在涵盖感知与推理任务的多个数据集上,对3B至72B规模的开源VLM及Gemini 2.0等专有模型进行了全面实验,控制变量包括样本数、检索方法、推理质量与分布差异。结果表明,模型对这些因素变化均不敏感,说明当前VLM未能有效利用演示信息以实现预期的多模态上下文学习。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are widely assumed to exhibit in-context learning (ICL), a property similar to that of their language-only counterparts. While recent work suggests VLMs can perform multimodal ICL (MM-ICL), studies show they often rely on shallow heuristics -- such as copying or majority voting -- rather than true task understanding. We revisit this assumption by evaluating VLMs under distribution shifts, where support examples come from a dataset different from the query. Surprisingly, performance often degrades with more demonstrations, and models tend to copy answers rather than learn from them. To investigate further, we propose a new MM-ICL with Reasoning pipeline that augments each demonstration with a generated rationale alongside the answer. We conduct extensive and comprehensive experiments on both perception- and reasoning-required datasets with open-source VLMs ranging from 3B to 72B and proprietary models such as Gemini 2.0. We conduct controlled studies varying shot count, retrieval method, rationale quality, and distribution. Our results show limited performance sensitivity across these factors, suggesting that current VLMs do not effectively utilize demonstration-level information as intended in MM-ICL.

多模态学习上下文学习视觉语言模型推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。