用专家推理轨迹指导模型,让视觉信息更好融入逻辑推理。
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
- 用强模型的推理轨迹引导弱模型,提升视觉证据使用
- 在多个多模态推理数据集上性能显著提升
- 适合需要精准视觉-逻辑对齐的研究者和开发者
近年来,基于可验证奖励的强化学习(RLVR)在多模态大语言模型(MLLMs)中的进展主要集中在提高最终答案正确性和强化视觉定位能力。然而,一个关键瓶颈依然存在:尽管模型能关注相关视觉区域,却常无法有效将视觉证据融入后续推理,导致推理链条与视觉事实结合较弱。为此,我们提出轨迹引导强化学习(TGRL),利用更强模型生成的专家推理轨迹,引导策略模型在细粒度推理过程中整合视觉证据。我们进一步引入词元级重加权和轨迹过滤机制,以确保策略优化的稳定性和有效性。在多个多模态推理基准上的大量实验表明,TGRL能持续提升推理性能,并有效弥合视觉感知与逻辑推理之间的差距。
原文摘要 · Abstract (English)
Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) for multimodal large language models (MLLMs) have mainly focused on improving final answer correctness and strengthening visual grounding. However, a critical bottleneck remains: although models can attend to relevant visual regions, they often fail to effectively incorporate visual evidence into subsequent reasoning, leading to reasoning chains that are weakly grounded in visual facts. To address this issue, we propose Trajectory-Guided Reinforcement Learning (TGRL), which guides the policy model to integrate visual evidence into fine-grained reasoning processes using expert reasoning trajectories from stronger models. We further introduce token-level reweighting and trajectory filtering to ensure stable and effective policy optimization. Extensive experiments on multiple multimodal reasoning benchmarks demonstrate that TGRL consistently improves reasoning performance and effectively bridges the gap between visual perception and logical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。