用视觉敏感度增强熵值,让强化学习更懂图像推理
Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection

- 将视觉敏感性与词元熵相乘耦合,精准定位关键信息
- 3B和7B模型上分别提升3.15和2.28分,显著超越基线
- 适合需要视觉-语义协同的智能体决策任务
尽管词元级熵在仅文本的强化学习中被广泛用于信用分配,但在视觉推理任务中其有效性尚未明确。我们通过控制实验发现,由于视觉敏感词元天然熵低,该机制在视觉推理中会失效。现有多模态强化学习方法虽重视视觉感知,却难以兼顾精确感知与语义推理,或缺乏系统性视觉度量,或忽视熵对语义探索的核心作用。为此,我们提出VEPO(视觉熵词元选择策略),通过原理性的乘法耦合,将视觉敏感性与词元熵结合,使梯度信用聚焦于同时具备视觉锚定和高信息量的词元。大量实验表明,VEPO在3B和7B规模下分别比仅使用熵的基线提升3.15和2.28分。消融实验进一步验证了方法的有效性。
原文摘要 · Abstract (English)
While token-level entropy is commonly recognized as effective for credit assignment in text-only reinforcement learning with verifiable rewards (RLVR), it remains unclear whether this mechanism still holds in visual reasoning. Our controlled study shows that this mechanism collapses in visual reasoning due to the omission of vision-sensitive tokens with naturally low entropy. Although existing multimodal RL methods increasingly acknowledge the importance of visual perception, they struggle to satisfy the inherent demand for interleaving precise perceptual grounding with semantic reasoning, either lacking systematic visual measurements or overlooking that token entropy primarily drives semantic exploration. To address this, we introduce VEPO (Vision-Entropy token-selection for Policy Optimization), an effective RL framework explicitly integrating visual sensitivity with token entropy via a principled multiplicative coupling, where VEPO redirects gradient credit toward tokens which are simultaneously visually grounded and highly informative. Extensive experiments demonstrate VEPO's leading performance, significantly outperforming the entropy-only baseline by 2.28 points at 7B-scale and 3.15 points at 3B-scale. Ablations further substantiate the soundness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。