arXiv:2604.18512cs.CV2026-04ACL

构建分层多图推理数据,提升视觉语言模型跨图理解能力

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models

论文配图:S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models
图 1 · 摘自论文原文
  • 设计从易到难的三阶推理框架,覆盖局部定位、跨图对比和全局搜索
  • 在LLaVA和Qwen-VL上显著提升多图推理性能,优于基线方法
  • 无需依赖特定模型特征,通用性强,适合希望增强多图理解的研究者

视觉语言模型在单图理解上已取得显著进展,但在多图推理方面仍存在挑战。现有方法主要聚焦于预设图像索引的局部推理(如'看第3张图'),忽视了全局视觉搜索和自主跨图比较等关键能力。为此,我们提出简单到困难(S2H)学习框架,系统构建三个层级的多图偏好数据:(1)单图局部推理,(2)多图局部对比,(3)全局视觉搜索。不同于依赖模型特定属性(如幻觉或注意力启发)生成偏好对的方法,本方案通过提示驱动的复杂度构造选中/拒绝对,适用于不同模型。在LLaVA和Qwen-VL上的广泛评估表明,所生成的多样化多图推理数据显著提升了多图推理表现,优于基线方法。重要的是,该方法在保持单图推理性能的同时,有效增强了多图理解能力,推动了整体视觉偏好对齐的前沿水平。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated remarkable progress in single-image understanding, yet effective reasoning across multiple images remains challenging. We identify a critical capability gap in existing multi-image alignment approaches: current methods focus primarily on localized reasoning with pre-specified image indices (``Look at Image 3 and...''), bypassing the essential skills of global visual search and autonomous cross-image comparison. To address this limitation, we introduce a Simple-to-Hard (S2H) learning framework that systematically constructs multi-image preference data across three hierarchical reasoning levels requiring an increasing level of capabilities: (1) single-image localized reasoning, (2) multi-image localized comparison, and (3) global visual search. Unlike prior work that relies on model-specific attributes, such as hallucinations or attention heuristics, to generate preference pairs, our approach leverages prompt-driven complexity to create chosen/rejected pairs that are applicable across different models. Through extensive evaluations on LLaVA and Qwen-VL models, we show that our diverse multi-image reasoning data significantly enhances multi-image reasoning performance, yielding significant improvements over baseline methods across benchmarks. Importantly, our approach maintains strong single-image reasoning performance while simultaneously strengthening multi-image understanding capabilities, thus advancing the state of the art for holistic visual preference alignment.

多图推理视觉语言模型偏好优化分层训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。