构建分层多图推理数据,提升视觉语言模型跨图理解能力
S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models

- 设计从易到难的三阶推理框架,覆盖局部定位、跨图对比和全局搜索
- 在LLaVA和Qwen-VL上显著提升多图推理性能,优于基线方法
- 无需依赖特定模型特征,通用性强,适合希望增强多图理解的研究者
视觉语言模型在单图理解上已取得显著进展,但在多图推理方面仍存在挑战。现有方法主要聚焦于预设图像索引的局部推理(如'看第3张图'),忽视了全局视觉搜索和自主跨图比较等关键能力。为此,我们提出简单到困难(S2H)学习框架,系统构建三个层级的多图偏好数据:(1)单图局部推理,(2)多图局部对比,(3)全局视觉搜索。不同于依赖模型特定属性(如幻觉或注意力启发)生成偏好对的方法,本方案通过提示驱动的复杂度构造选中/拒绝对,适用于不同模型。在LLaVA和Qwen-VL上的广泛评估表明,所生成的多样化多图推理数据显著提升了多图推理表现,优于基线方法。重要的是,该方法在保持单图推理性能的同时,有效增强了多图理解能力,推动了整体视觉偏好对齐的前沿水平。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have demonstrated remarkable progress in single-image understanding, yet effective reasoning across multiple images remains challenging. We identify a critical capability gap in existing multi-image alignment approaches: current methods focus primarily on localized reasoning with pre-specified image indices (``Look at Image 3 and...''), bypassing the essential skills of global visual search and autonomous cross-image comparison. To address this limitation, we introduce a Simple-to-Hard (S2H) learning framework that systematically constructs multi-image preference data across three hierarchical reasoning levels requiring an increasing level of capabilities: (1) single-image localized reasoning, (2) multi-image localized comparison, and (3) global visual search. Unlike prior work that relies on model-specific attributes, such as hallucinations or attention heuristics, to generate preference pairs, our approach leverages prompt-driven complexity to create chosen/rejected pairs that are applicable across different models. Through extensive evaluations on LLaVA and Qwen-VL models, we show that our diverse multi-image reasoning data significantly enhances multi-image reasoning performance, yielding significant improvements over baseline methods across benchmarks. Importantly, our approach maintains strong single-image reasoning performance while simultaneously strengthening multi-image understanding capabilities, thus advancing the state of the art for holistic visual preference alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。