arXiv:2504.16656cs.CV2025-04被引 43

混合强化学习提升多模态推理能力,性能逼近顶级闭源模型。

Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

  • 融合MPO与GRPO的混合强化学习框架,兼顾推理深度与泛化能力。
  • 在OlympiadBench等4个基准上表现领先,最高达78.9分。
  • 开源380亿参数模型,适合研究多模态推理与开放复现者使用。

我们提出Skywork R1V2,下一代多模态推理模型,相较于前代Skywork R1V实现重大突破。其核心是混合强化学习范式,联合使用混合偏好优化(MPO)与组相对策略优化(GRPO),协调奖励模型引导与规则策略,有效解决复杂推理能力与广泛泛化之间的长期难题。为提升训练效率,提出选择性样本缓冲(SSB)机制,通过优先处理高价值样本缓解GRPO中的优势消失问题。值得注意的是,过度强化信号会引发视觉幻觉,我们通过校准奖励阈值系统性监测并抑制该现象。实证结果表明,R1V2表现卓越:在OlympiadBench上达62.6,在AIME2024上达78.9,在LiveCodeBench上达63.6,在MMMU上达73.6。这些结果证明其优于现有开源模型,显著缩小与顶级闭源系统(如Gemini 2.5和OpenAI-o4-mini)的差距。模型权重已公开,促进开放与可复现性,地址:https://huggingface.co/Skywork/Skywork-R1V2-38B。

原文摘要 · Abstract (English)

We present Skywork R1V2, a next-generation multimodal reasoning model and a major leap forward from its predecessor, Skywork R1V. At its core, R1V2 introduces a hybrid reinforcement learning paradigm that jointly leverages the Mixed Preference Optimization (MPO) and the Group Relative Policy Optimization (GRPO), which harmonizes reward-model guidance with rule-based strategies, thereby addressing the long-standing challenge of balancing sophisticated reasoning capabilities with broad generalization. To further enhance training efficiency, we propose the Selective Sample Buffer (SSB) mechanism, which effectively addresses the vanishing advantages dilemma inherent in GRPO by prioritizing high-value samples throughout the optimization process. Notably, we observe that excessive reinforcement signals can induce visual hallucinations--a phenomenon we systematically monitor and mitigate through calibrated reward thresholds throughout the training process. Empirical results affirm the exceptional capability of R1V2, with benchmark-leading performances such as 62.6 on OlympiadBench, 78.9 on AIME2024, 63.6 on LiveCodeBench, and 73.6 on MMMU. These results underscore R1V2's superiority over existing open-source models and demonstrate significant progress in closing the performance gap with premier proprietary systems, including Gemini 2.5 and OpenAI-o4-mini. The Skywork R1V2 model weights have been publicly released to promote openness and reproducibility https://huggingface.co/Skywork/Skywork-R1V2-38B.

多模态推理强化学习开源模型混合优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。