arXiv:2605.23997cs.CVcs.AI2026-05

通过迭代视觉对齐修正推理路径,提升长时程多模态任务的逻辑与视觉一致性。

IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning

论文配图:IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
图 1 · 摘自论文原文
  • 引入迭代视觉对齐机制,动态修正推理过程中的错误轨迹。
  • 在多个多模态基准上超越现有强化学习方法,准确率显著提升。
  • 适合需要高精度视觉与逻辑一致性的复杂多模态任务研究者。

基于强化学习的多模态大语言模型在复杂视觉推理任务中表现出色,但在长时程多模态场景下仍受限于视觉幻觉和逻辑错误。现有方法通常将高维视觉场景预编码为离散文本代理以支持下游推理,但随着推理链展开,文本与视觉之间的信息不对称会削弱视觉接地性,导致误导性推理和错误输出。为此,我们提出IVR-R1(迭代视觉对齐推理),一种新型强化学习训练框架,通过奖励驱动的筛选机制识别有缺陷的轨迹,并在多模态上下文中进行细粒度的步骤级错误归因。通过迭代交叉验证中间推理状态与原始视觉先验,重建推理循环可自动修正推理路径,生成高质量的专家级推理模板,用于指导策略模型优化。在多个多模态基准上的实验表明,IVR-R1持续优于现有强化学习方法,建立了一种维持复杂多模态推理中逻辑与视觉一致性的更优范式。

原文摘要 · Abstract (English)

Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they remain limited in long-horizon multimodal scenarios, often suffering from visual hallucination and logical error. Current methods typically pre-encode high-dimensional visual scenes into discrete textual proxies to facilitate downstream reasoning. As the reasoning chain unfolds, however, the inherent information asymmetry between text and visual scenes tends to erode visual grounding, resulting in misguided reasoning and erroneous outputs. To address this issue, we introduce IVR-R1 (Iterative Visual-grounded Reasoning), a novel RL training framework that facilitates dynamic visual re-alignment that actively rectifies reasoning trajectories to guide policy optimization. Specifically, by leveraging a reward-driven screening mechanism to identify flawed rollouts, IVR-R1 executes a fine-grained, step-level error attribution within the multimodal context. By iteratively cross-referencing intermediate reasoning states against pristine visual priors, a Re-Reasoning Loop enables automated trajectory rectification, effectively synthesizing expert-level demonstrations that serve as high-fidelity reasoning templates for the policy model. Our experiments across diverse multimodal benchmarks demonstrate that IVR-R1 consistently outperforms existing reinforcement learning methods, establishing a superior paradigm for maintaining logical and visual consistency in complex multimodal reasoning.

强化学习多模态推理视觉对齐轨迹修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。