让视觉模型更懂图像关键区域,提升多模态推理能力
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
- 通过分层几何聚合定位图像重要区域
- 在多个基准上提升各类统一奖励方法性能
- 适合需要精准视觉理解的多模态任务研究者
强化学习从可验证奖励(RLVR)显著提升了大语言模型在抽象推理任务中的表现。然而,其在大视觉语言模型(LVLMs)上的应用仍受限于结构化表征瓶颈。现有方法普遍缺乏对视觉信息的显式建模与有效利用,导致视觉表示难以与强化学习优化过程紧密耦合,限制了多模态推理性能的进一步提升。为此,我们提出KAWHI(关键区域对齐加权谐振激励),一种即插即用的奖励重加权机制,将结构化视觉信息显式融入统一奖励策略优化方法(如GRPO和GSPO)。该方法通过分层几何聚合自适应定位语义显著区域,利用结构化归因识别视觉关键注意力头,并进行段落级信用再分配,使空间视觉证据与语义决定性推理步骤对齐。在多种推理基准上的大量实证评估表明,KAWHI作为通用增强模块,能持续提升各类统一奖励优化方法的性能。
原文摘要 · Abstract (English)
Reinforcement Learning from Verifiable Rewards (RLVR) has substantially enhanced the reasoning capabilities of large language models in abstract reasoning tasks. However, its application to Large Vision-Language Models (LVLMs) remains constrained by a structural representational bottleneck. Existing approaches generally lack explicit modeling and effective utilization of visual information, preventing visual representations from being tightly coupled with the reinforcement learning optimization process and thereby limiting further improvements in multimodal reasoning performance. To address this limitation, we propose KAWHI (Key-Region Aligned Weighted Harmonic Incentive), a plug-and-play reward reweighting mechanism that explicitly incorporates structured visual information into uniform reward policy optimization methods (e.g., GRPO and GSPO). The method adaptively localizes semantically salient regions through hierarchical geometric aggregation, identifies vision-critical attention heads via structured attribution, and performs paragraph-level credit reallocation to align spatial visual evidence with semantically decisive reasoning steps. Extensive empirical evaluations on diverse reasoning benchmarks substantiate KAWHI as a general-purpose enhancement module, consistently improving the performance of various uniform reward optimization methods. Project page: KAWHI (https://kawhiiiileo.github.io/KAWHI_PAGE/)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。