通过视觉焦点链提升多图理解能力,让AI更懂复杂图像场景。
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
- 构建焦点中心视觉链,串联多图关键信息
- 在7个基准上平均提升3.16%和2.24%
- 适合需要跨图推理的视觉语言任务
视觉语言模型(VLMs)在单图任务中表现卓越,但在涉及复杂多图输入的实际场景中,性能显著下降,因模型难以从分散的视觉特征中提取关键信息。本文提出焦点中心视觉链(Focus-Centric Visual Chain)新范式,增强VLM在多图场景中的感知、理解与推理能力。为此,我们设计了可扩展的自底向上数据合成方法——焦点中心数据合成(Focus-Centric Data Synthesis),构建了大规模数据集VISC-150K,包含以焦点中心视觉链形式呈现的复杂推理路径,专为多图任务设计。在7个多图基准上的实验表明,该方法在两种不同模型架构上分别实现平均3.16%和2.24%的性能提升,且不损害通用视觉语言能力。本研究推动了更鲁棒、更强的视觉语言系统向复杂视觉场景的迈进。
原文摘要 · Abstract (English)
Vision-language models (VLMs) achieve remarkable success in single-image tasks. However, real-world scenarios often involve intricate multi-image inputs, leading to a notable performance decline as models struggle to disentangle critical information scattered across complex visual features. In this work, we propose Focus-Centric Visual Chain, a novel paradigm that enhances VLMs'perception, comprehension, and reasoning abilities in multi-image scenarios. To facilitate this paradigm, we propose Focus-Centric Data Synthesis, a scalable bottom-up approach for synthesizing high-quality data with elaborate reasoning paths. Through this approach, We construct VISC-150K, a large-scale dataset with reasoning data in the form of Focus-Centric Visual Chain, specifically designed for multi-image tasks. Experimental results on seven multi-image benchmarks demonstrate that our method achieves average performance gains of 3.16% and 2.24% across two distinct model architectures, without compromising the general vision-language capabilities. our study represents a significant step toward more robust and capable vision-language systems that can handle complex visual scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。