arXiv:2507.03019cs.CVcs.LG2025-07AAAI被引 62

让大模型自己学会回头看图像,提升推理能力

Look-Back: Implicit Visual Re-focusing in MLLM Reasoning

  • 通过分析注意力模式,发现模型能自发关注视觉信息
  • 无需额外输入或结构修改,显著提升多模态推理表现
  • 适合需要精准视觉理解的智能问答与图像推理场景

多模态大语言模型在多模态推理任务中取得了显著进展,但在推理后期常过度依赖文本信息,忽视视觉输入。现有方法通常通过显式注入视觉信息来缓解此问题。本文通过对MLLM注意力模式的分析,发现:在适当引导下,模型能在推理后期自发重新聚焦于视觉信息,无需显式注入。这一现象表明MLLM具备内在的视觉融合推理能力。基于此,我们提出Look-Back——一种隐式引导机制,使模型在推理过程中自主决定何时、何地、如何“回看”视觉内容。该方法无需改变模型结构或添加输入,经多个多模态基准测试验证,显著提升了模型的推理与感知能力。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in multimodal reasoning. However, they often excessively rely on textual information during the later stages of inference, neglecting the crucial integration of visual input. Current methods typically address this by explicitly injecting visual information to guide the reasoning process. In this work, through an analysis of MLLM attention patterns, we made an intriguing observation: with appropriate guidance, MLLMs can spontaneously re-focus their attention on visual inputs during the later stages of reasoning, even without explicit visual information injection. This spontaneous shift in focus suggests that MLLMs are intrinsically capable of performing visual fusion reasoning. Building on this insight, we introduce Look-Back, an implicit approach designed to guide MLLMs to ``look back" at visual information in a self-directed manner during reasoning. Look-Back empowers the model to autonomously determine when, where, and how to re-focus on visual inputs, eliminating the need for explicit model-structure constraints or additional input. We demonstrate that Look-Back significantly enhances the model's reasoning and perception capabilities, as evidenced by extensive empirical evaluations on multiple multimodal benchmarks.

多模态推理注意力机制视觉回溯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。