arXiv:2606.08708cs.CV2026-06被引 2

通过细粒度奖励重塑,让视觉关键词获得更强监督。

PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping

论文配图:PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping
图 1 · 摘自论文原文
  • 按视觉依赖性筛选关键视觉词,动态调整每步奖励
  • 在7个基准上提升21%-23%,模型规模越大效果越明显
  • 适合需要精准视觉推理的多模态任务研究者

基于可验证奖励的强化学习(RLVR)已成为提升大视觉语言模型(LVLMs)推理能力的有效范式。然而,现有方法主要依赖轨迹级结果奖励,对所有生成标记分配相同学习信号,这种粗粒度信用分配与多模态推理不匹配——仅少数标记与视觉证据因果相关。为此,我们提出感知增强策略优化(PRPO),一种细粒度的令牌级强化学习框架,显式识别并强化长时程多模态推理轨迹中的关键感知标记。PRPO引入鲁棒视觉依赖性(RVD)度量,识别既具视觉依据又对扰动稳定的标记,过滤掉脆弱或噪声标记。基于RVD,进一步提出感知优势重塑(PAR),一种令牌级信用分配技术,放大具有感知信息的标记奖励,同时保持非感知标记的稳定梯度。在七个多模态推理基准上的实验表明,PRPO在3B和7B模型规模下均显著优于强基线,平均提升分别为23.3%和21.1%。该方法实现当前最优性能,训练效率更高,跨任务泛化能力更强。研究凸显了细粒度信用分配在可扩展多模态强化学习中的重要性。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective paradigm for improving the reasoning capability of Large Vision-Language Models (LVLMs). However, existing RLVR methods primarily rely on trajectory-level outcome rewards, which assign identical learning signals across all generated tokens. This coarse-grained credit assignment is fundamentally mismatched to multimodal reasoning, where only a sparse subset of tokens is causally grounded in visual evidence. Consequently, these pivotal perceptual tokens receive weak supervision and are often overwhelmed by language priors or reasoning-template tokens. To address this limitation, we propose Perception-Reinforced Policy Optimization (PRPO), a token-level reinforcement learning framework that explicitly identifies and reinforces pivotal perceptual tokens within long-horizon multimodal reasoning trajectories. PRPO introduces Robust Visual Dependency (RVD), a principled metric that identifies tokens whose predictions are both visually grounded and perturbation-stable, filtering out brittle or noisy visual tokens. Based on RVD, we further propose Perceptual Advantage Reshaping (PAR), a token-level credit assignment technique that amplifies perceptually informative tokens while preserving stable gradients for non-perceptual tokens. Extensive experiments on seven multimodal reasoning benchmarks demonstrate that PRPO consistently outperforms strong LVLM baselines across both 3B and 7B model scales, achieving average gains of 23.3% and 21.1%, respectively. PRPO achieves state-of-the-art performance with improved training efficiency and stronger cross-task generalization. Our findings highlight the importance of fine-grained credit assignment for scalable multimodal reinforcement learning.

多模态推理强化学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。