arXiv:2603.25077cs.CV2026-03被引 6

让多模态大模型同时精准理解视觉和逻辑推理

Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs

  • 通过动态重加权感知与推理类标记,显式建模两者耦合关系
  • 在多个多模态推理数据集上实现领先性能,视觉与逻辑均准确
  • 可插拔适配现有强化学习方法,提升稳定性和泛化能力

将可验证奖励的强化学习(RLVR)拓展至多模态大语言模型(MLLMs)面临根本挑战:其输出天然混合了与视觉内容相关的感知标记和构建推理链的符号推理标记。这两类标记分别体现视觉定位与符号推理能力,但相互依赖,孤立优化效果不佳。通过标记级实证分析,我们发现仅优化感知或推理标记均表现逊于全量优化,凸显其内在耦合性。为此,提出即插即用的标记重加权(ToR)策略,通过识别两类关键标记并动态调整权重,显式建模这种依赖关系。在现有方法(如GRPO和DAPO)基础上应用ToR,显著提升多个多模态推理基准的表现,实现最优性能,兼具准确视觉定位与连贯推理。

原文摘要 · Abstract (English)

Extending Reinforcement Learning with Verifiable Rewards (RLVR) to multimodal large language models (MLLMs) faces a fundamental challenge: their responses inherently interleave perception-related tokens, which ground visual content, with reasoning-related tokens, which construct reasoning chains. These token types instantiate distinct yet interdependent capacities -- visual grounding and symbolic reasoning -- making isolated optimization insufficient. Through token-level empirical analysis, we demonstrate that optimizing either perception- or reasoning-only tokens consistently underperforms full optimization, underscoring their inherent coupling. To address this, we propose a plug-and-play Token-Reweighting (ToR) strategy that explicitly models this interdependence by identifying critical tokens of both types and dynamically reweighting them during RLVR training. Applied on top of existing methods (e.g., GRPO and DAPO), ToR delivers consistent performance gains across multiple multi-modal reasoning benchmarks, achieving state-of-the-art performance with both accurate visual grounding and coherent reasoning.

多模态强化学习推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。