arXiv:2601.07812cs.CV2026-01ACL被引 6

揭示视觉语言模型多图理解的三大缺陷并提出改进方案

More Images, More Problems? A Controlled Analysis of VLM Failure Modes

  • 构建新基准MIMIC,系统诊断多图模型能力短板
  • 多图信息整合能力提升,跨图任务性能超越当前最佳
  • 适合研究多模态模型缺陷与训练优化的学者

大型视觉语言模型(LVLM)在单图理解上表现卓越,但在多图理解和推理方面仍缺乏深入研究。现有评估基准虽已开始测试多图模型,但对其核心弱点及成因的系统分析仍不足。本文提出MIMIC(多图模型洞察与挑战)基准,用于严格评估LVLM的多图能力。通过一系列诊断实验,发现LVLM普遍存在跨图信息整合失败、难以同时跟踪多个概念的问题。为此,我们提出两种互补性解决方案:数据层面,设计一种过程式生成策略,将单图标注组合为结构化多图训练样本;优化层面,分析分层注意力模式,提出针对多图输入的注意力掩码机制。实验表明,该方法显著提升了跨图信息聚合能力,并在现有多个多图基准上超越先前最优表现。数据与代码将公开于https://github.com/anurag-198/MIMIC。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) have demonstrated remarkable capabilities, yet their proficiency in understanding and reasoning over multiple images remains largely unexplored. While existing benchmarks have initiated the evaluation of multi-image models, a comprehensive analysis of their core weaknesses and their causes is still lacking. In this work, we introduce MIMIC (Multi-Image Model Insights and Challenges), a new benchmark designed to rigorously evaluate the multi-image capabilities of LVLMs. Using MIMIC, we conduct a series of diagnostic experiments that reveal pervasive issues: LVLMs often fail to aggregate information across images and struggle to track or attend to multiple concepts simultaneously. To address these failures, we propose two novel complementary remedies. On the data side, we present a procedural data-generation strategy that composes single-image annotations into rich, targeted multi-image training examples. On the optimization side, we analyze layer-wise attention patterns and derive an attention-masking scheme tailored for multi-image inputs. Experiments substantially improved cross-image aggregation, while also enhancing performance on existing multi-image benchmarks, outperforming prior state of the art across tasks. Data and code will be made available at https://github.com/anurag-198/MIMIC.

多图理解视觉语言模型模型缺陷分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。