arXiv:2507.06167cs.CLcs.CV2025-07被引 11

用强化学习让开源视觉语言模型学会推理,性能逼近人类水平。

Skywork-R1V3 Technical Report

  • 通过强化学习后训练框架激活模型推理能力,无需额外预训练。
  • 在MMMU数据集上准确率从64.3%提升至76.0%,达到入门级人类水平。
  • 适用于希望提升开源多模态模型推理能力的研究者与开发者。

我们介绍Skywork-R1V3,一个先进的开源视觉语言模型(VLM),开创性地将纯文本大语言模型(LLM)的推理能力迁移至视觉任务。其核心创新在于通过精心设计的强化学习后训练框架,有效激活并增强模型的推理能力,无需额外继续预训练。该框架揭示了连接模块在实现跨模态对齐中的关键作用。此外,我们提出一种新的推理能力指标——关键推理标记的熵值,在强化学习训练中用于检查点选择,效果显著。Skywork-R1V3在MMMU数据集上表现优异,准确率从64.3%提升至76.0%,达到入门级人类水平。令人瞩目的是,该强化学习方法使38B参数模型也能媲美顶尖闭源VLM。模型成功将数学推理能力迁移至其他学科相关任务。论文还分析了课程学习与强化微调策略,并展开对多模态推理的广泛讨论。Skywork-R1V3标志着多模态推理的重大进展,彰显强化学习在推动开源VLM发展中的强大潜力。

原文摘要 · Abstract (English)

We introduce Skywork-R1V3, an advanced, open-source vision-language model (VLM) that pioneers a new approach to visual reasoning. Its key innovation lies in effectively transferring reasoning skills from text-only Large Language Models (LLMs) to visual tasks. The strong performance of Skywork-R1V3 primarily stems from our elaborate post-training RL framework, which effectively activates and enhances the model's reasoning ability, without the need for additional continue pre-training. Through this framework, we further uncover the fundamental role of the connector module in achieving robust cross-modal alignment for multimodal reasoning models. In addition, we introduce a unique indicator of reasoning capability, the entropy of critical reasoning tokens, which has proven highly effective for checkpoint selection during RL training. Skywork-R1V3 achieves state-of-the-art results on MMMU, significantly improving from 64.3% to 76.0%. This performance matches entry-level human capabilities. Remarkably, our RL-powered post-training approach enables even the 38B parameter model to rival top closed-source VLMs. The implementation successfully transfers mathematical reasoning to other subject-related reasoning tasks. We also include an analysis of curriculum learning and reinforcement finetuning strategies, along with a broader discussion on multimodal reasoning. Skywork-R1V3 represents a significant leap in multimodal reasoning, showcasing RL as a powerful engine for advancing open-source VLM capabilities.

视觉语言模型多模态推理强化学习开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。