arXiv:2510.09285cs.CV2025-10中稿 · ICLR被引 37

通过分析生成词元的视觉依赖性,提升多模态强化学习的推理能力。

Spotlight on Token Perception for Multimodal Reinforcement Learning

  • 从词元级视角衡量视觉依赖性,发现视觉相关词元稀疏分布。
  • 提出VPPO算法,在8个基准上优于主流开源模型,7B和32B均有效。
  • 适合关注多模态推理优化与视觉感知建模的研究者。

尽管可验证奖励强化学习(RLVR)提升了大视觉语言模型(LVLMs)的推理能力,但现有方法在多模态推理中常忽略视觉感知在RLVR优化中的关键作用。本文首次从词元感知角度探索多模态RLVR,通过细粒度分析思维链(CoT)过程,发现:第一,轨迹中仅少数词元具有高视觉依赖性;第二,不同轨迹间整体视觉依赖性差异显著。基于此,提出视觉感知策略优化(VPPO)算法,采用双重机制:用整体视觉依赖性重加权轨迹优势,并仅对感知关键词元进行策略更新。在涵盖8个感知与推理任务的综合基准测试中,VPPO在7B与32B模型规模下均显著优于领先开源模型,效果持续验证。研究不仅建立了词元级视觉感知分析新视角,还提出一种高效优化策略,显著增强LVLM的多模态推理能力。

原文摘要 · Abstract (English)

While Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Vision-Language Models (LVLMs), most existing methods in multimodal reasoning neglect the critical role of visual perception within the RLVR optimization process. In this paper, we undertake a pioneering exploration of multimodal RLVR through the novel perspective of token perception, which measures the visual dependency of each generated token. With a granular analysis of Chain-of-Thought (CoT) processes, we uncover two key insights: first, token perception in a rollout trajectory is sparsely distributed, where only a small fraction of tokens have high visual dependency for visually-grounded reasoning; second, different trajectories exhibit significant divergence in their overall visual dependency. Based on these observations, we propose Visually-Perceptive Policy Optimization (VPPO), a novel policy gradient algorithm that explicitly leverages token perception to refine the learning signal. Specifically, VPPO achieves this through a dual mechanism: it reweights a trajectory's advantage by its overall visual dependency, and focuses policy updates exclusively on perceptually pivotal tokens. On a comprehensive suite of eight perception and reasoning benchmarks, VPPO demonstrates substantial gains over leading open-source RL-tuned models, with its effectiveness consistently validated across 7B and 32B model scales. Our findings not only establish a new token-level perceptual perspective for analyzing multimodal RLVR but also present a novel and effective optimization strategy to significantly enhance the multimodal reasoning capabilities of LVLMs.

多模态强化学习视觉感知推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。