arXiv:2605.22072cs.CLcs.CV2026-05被引 1

通过锚定与强化视觉注意力,提升多模态模型推理的可信度。

Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention

论文配图:Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
图 1 · 摘自论文原文
  • 用专门标记显式监督图像区域注意力,而非依赖文本描述。
  • 在小数据下优于现有基线,在多个基准上实现显著提升。
  • 适合关注多模态推理可靠性与视觉证据使用的研究者。

基于可验证奖励的强化学习(RLVR)为大语言模型的复杂推理提供了新范式,近期工作已将其拓展至多模态大语言模型(MLLMs)。然而,该迁移暴露了可信性挑战:任务相关视觉证据的准确感知与推理中对证据的忠实利用。现有感知监督常基于文本描述而非图像区域,且对证据的忠实使用未受重视,导致感知与推理脱节——正确感知的证据在推理中被忽略或矛盾。为此,我们提出 Faithful-MR1 训练框架,通过锚定与强化视觉注意力解决可信多模态推理的双重问题。锚定阶段将感知作为预推理子任务,直接监督专用 <Focus> 标记对图像区域的注意力;强化阶段通过反事实图像干预,奖励在视觉因果关键处集中注意力且答案正确的推理轨迹。大量实验表明,Faithful-MR1 在 Qwen2.5-VL-Instruct 3B 与 7B 模型上均超越近期多模态推理基线,且训练数据使用量显著更少。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for advancing complex reasoning in large language models, and recent work extends RLVR to multimodal large language models (MLLMs). This transfer, however, surfaces a faithfulness challenge: faithful perception of task-relevant visual evidence and faithful use of that evidence during reasoning, leading to unsatisfactory gains on multimodal benchmarks. Specifically, existing perception supervision often operates on textual descriptions rather than natively on image regions, and faithful use is largely overlooked, exposing the perception-reasoning disconnect where correctly perceived evidence is dropped or contradicted during reasoning. To close these gaps, we propose Faithful-MR1, a training framework that anchors and reinforces visual attention to address both halves of faithful multimodal reasoning. The Anchoring stage turns perception into an explicit pre-reasoning subtask, supervising a dedicated <Focus> token's attention directly against image regions rather than through textual descriptions. The Reinforcing stage exposes faithful use through counterfactual image intervention, rewarding answer-correct trajectories that concentrate visual attention where vision causally matters. Extensive experiments demonstrate that Faithful-MR1 outperforms recent multimodal reasoning baselines on both Qwen2.5-VL-Instruct 3B and 7B backbones while using substantially less training data.

多模态推理视觉注意力强化学习可信性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。