提出新强化学习方法,让多模态模型更稳定地处理视觉推理任务。
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks
- 用非线性分布匹配替代线性缩放,实现跨任务梯度平衡
- 在18个基准上超越主流开源与闭源模型,性能领先
- 适合需要强视觉理解与复杂推理的通用多模态应用
Group Relative Policy Optimization(GRPO)已成为驱动多模态大语言模型进步的主流强化学习目标。然而,将这一成功拓展至开源多模态通用模型仍受两大挑战制约:不同视觉任务间奖励拓扑的极端差异,以及精细感知与多步推理能力之间的难以平衡。为此,我们提出高斯GRPO(G²RPO),一种新型强化学习训练目标,以非线性分布匹配取代标准线性缩放。通过数学上强制任意任务的优势分布严格收敛至标准正态分布𝒩(0,1),G²RPO理论上确保了跨任务梯度公平性,缓解重尾异常值的脆弱性,并对正负奖励提供对称更新。基于G²RPO带来的增强训练稳定性,我们引入两种任务级奖励塑造机制,无缝平衡感知与推理能力。首先,响应长度塑造动态激发复杂查询的长推理链,同时强制直接输出以强化视觉定位;其次,熵塑造严格约束模型探索范围,有效防止熵坍缩与熵爆炸。结合上述方法,我们提出OpenVLThinkerV2,一个高度鲁棒、通用的多模态模型。在18个多样化基准上的广泛评估表明,其性能显著优于强大开源模型及领先的专有前沿模型。
原文摘要 · Abstract (English)
Group Relative Policy Optimization (GRPO) has emerged as the de facto Reinforcement Learning (RL) objective driving recent advancements in Multimodal Large Language Models. However, extending this success to open-source multimodal generalist models remains heavily constrained by two primary challenges: the extreme variance in reward topologies across diverse visual tasks, and the inherent difficulty of balancing fine-grained perception with multi-step reasoning capabilities. To address these issues, we introduce Gaussian GRPO (G$^2$RPO), a novel RL training objective that replaces standard linear scaling with non-linear distributional matching. By mathematically forcing the advantage distribution of any given task to strictly converge to a standard normal distribution, $\mathcal{N}(0,1)$, G$^2$RPO theoretically ensures inter-task gradient equity, mitigates vulnerabilities to heavy-tail outliers, and offers symmetric update for positive and negative rewards. Leveraging the enhanced training stability provided by G$^2$RPO, we introduce two task-level shaping mechanisms to seamlessly balance perception and reasoning. First, response length shaping dynamically elicits extended reasoning chains for complex queries while enforce direct outputs to bolster visual grounding. Second, entropy shaping tightly bounds the model's exploration zone, effectively preventing both entropy collapse and entropy explosion. Integrating these methodologies, we present OpenVLThinkerV2, a highly robust, general-purpose multimodal model. Extensive evaluations across 18 diverse benchmarks demonstrate its superior performance over strong open-source and leading proprietary frontier models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。