arXiv:2602.08241cs.AIcs.CV2026-02被引 4

用强化学习让多模态模型更准地看图,减少推理错误。

Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs

  • 用区域级视觉注意力奖励机制,引导模型聚焦关键图像区域。
  • 在多个基准上提升感知与推理任务表现,错误率显著降低。
  • 适合关注视觉注意力优化与多模态推理的 researchers。

尽管链式思维(CoT)推理已显著提升多模态大模型在复杂任务上的表现,但现有方法主要依赖长文本推理轨迹,缺乏稳定视觉注意力策略的学习机制。我们的分析表明,当前多模态大模型存在视觉焦点薄弱问题:早期视觉错位在后续推理中很少被纠正,导致错误传播和推断失败。我们认为这一局限源于训练中对视觉注意力的信用分配不足。为此,我们提出SAYO,一种基于强化学习框架的视觉推理模型,引入基于区域级视觉注意力的奖励机制。该奖励将优化信号显式对齐于视觉依据的推理步骤,使模型学会更可靠的注意力行为。在多个多模态基准上的广泛实验表明,SAYO在多样化的推理与感知任务上均持续提升性能。

原文摘要 · Abstract (English)

While chain-of-thought (CoT) reasoning has substantially improved multimodal large language models (MLLMs) on complex reasoning tasks, existing approaches largely rely on long textual reasoning trajectories and provide limited mechanisms for learning stable visual attention policies. Our analysis shows that current MLLMs exhibit weak visual focus: early-stage visual misalignment is rarely corrected during subsequent reasoning, leading to error propagation and failed inferences. We argue that this limitation stems from inadequate credit assignment for visual attention during training. To address this issue, we propose SAYO, a visual reasoning model trained with a reinforcement learning (RL) framework that introduces a region-level visual attention-based reward. This reward explicitly aligns optimization signals with visually grounded reasoning steps, enabling the model to learn more reliable attention behaviors. Extensive experiments across multiple multimodal benchmarks demonstrate that SAYO consistently improves performance on diverse reasoning and perception tasks.

多模态视觉注意力强化学习推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。