arXiv:2603.03857cs.CV2026-03被引 4

无需训练即可提升大模型视觉推理能力,通过分层扫描找回关键视觉线索。

DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models

  • 分层扫描+多尺度提取,逐级找回被干扰的视觉证据
  • 集成Qwen2.5-VL-7B时在V*数据集达90.6%准确率
  • 适用于各类模型架构,无需额外调参,适合部署到实际系统

人类在嘈杂环境中仍能通过识别关键线索并将其与整体上下文关联,实现鲁棒的视觉定位与有根据的回答。受此启发,我们提出DeepScan,一种无需训练的框架,结合分层扫描、重聚焦和证据增强推理,用于大视觉语言模型(LVLMs)的视觉接地推理。与现有方法一次性定位完整证据不同,分层扫描采用自底向上的方式探索局部线索并提取多尺度证据,有效缓解干扰上下文的影响。重聚焦通过LVLM与视觉专家协作优化定位视图。最后,证据增强推理利用混合证据记忆聚合多粒度视角,生成准确且可解释的答案。实验表明,DeepScan显著提升了多种视觉任务中LVLM的表现,尤其在细粒度视觉理解上效果突出。集成Qwen2.5-VL-7B后,在V*数据集上达到90.6%的整体准确率。此外,DeepScan在不同架构与规模的模型上均提供一致改进,且无需额外适应成本。

原文摘要 · Abstract (English)

Humans can robustly localize visual evidence and provide grounded answers even in noisy environments by identifying critical cues and then relating them to the full context in a bottom-up manner. Inspired by this, we propose DeepScan, a training-free framework that combines Hierarchical Scanning, Refocusing, and Evidence-Enhanced Reasoning for visually grounded reasoning in Large Vision-Language Models (LVLMs). Unlike existing methods that pursue one-shot localization of complete evidence, Hierarchical Scanning performs local cue exploration and multi-scale evidence extraction to recover evidence in a bottom-up manner, effectively mitigating the impacts of distractive context. Refocusing then optimizes the localized evidence view through collaboration of LVLMs and visual experts. Finally, Evidence-Enhanced Reasoning aggregates multi-granular views via a hybrid evidence memory and yields accurate and interpretable answers. Experimental results demonstrate that DeepScan significantly boosts LVLMs in diverse visual tasks, especially in fine-grained visual understanding. It achieves 90.6% overall accuracy on V* when integrated with Qwen2.5-VL-7B. Moreover, DeepScan provides consistent improvements for LVLMs across various architectures and model scales without additional adaptation cost.

视觉推理无训练框架多尺度提取大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。