arXiv:2505.23766cs.CV2025-05CVPR被引 49

让AI看图推理更精准,通过聚焦关键物体提升视觉理解能力

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

  • 用物体为中心的视觉提示作为思维链信号,引导注意力
  • 在多个基准测试中显著提升图像推理与指代定位准确率
  • 适合需要精细视觉分析的应用,如医疗影像、自动驾驶

多模态大模型在视觉-语言任务中表现卓越,但在需精准视觉关注的场景下仍存在困难。本文提出Argus,引入一种新的视觉注意力锚定机制,以物体为中心的接地信号作为视觉思维链,实现目标条件下的有效视觉注意力。在多个基准上的评估表明,Argus在多模态推理和指代物体定位任务中均表现优异。深入分析验证了其设计选择的有效性,揭示了显式语言引导的视觉兴趣区域参与对多模态模型的重要性,强调从视觉中心视角推进多模态智能发展的必要性。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks, yet they often struggle with vision-centric scenarios where precise visual focus is needed for accurate reasoning. In this paper, we introduce Argus to address these limitations with a new visual attention grounding mechanism. Our approach employs object-centric grounding as visual chain-of-thought signals, enabling more effective goal-conditioned visual attention during multimodal reasoning tasks. Evaluations on diverse benchmarks demonstrate that Argus excels in both multimodal reasoning tasks and referring object grounding tasks. Extensive analysis further validates various design choices of Argus, and reveals the effectiveness of explicit language-guided visual region-of-interest engagement in MLLMs, highlighting the importance of advancing multimodal intelligence from a visual-centric perspective. Project page: https://yunzeman.github.io/argus/

视觉推理多模态思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。