VAPO让推理模型在5000步内稳定达到60.4分,性能超前人10分以上。
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
- 基于价值函数设计新框架,解决长思维链中的三大挑战
- 在AIME 2024上达60.4分,5000步内收敛且无训练崩溃
- 适合需要高效可靠推理的复杂任务研究者
我们提出VAPO(基于价值的增强近端策略优化框架),专为推理模型设计的价值基强化学习框架。在AIME 2024数据集上,基于Qwen 32B预训练模型的VAPO取得60.4的领先分数。在相同实验条件下,其性能显著优于DeepSeek-R1-Zero-Qwen-32B和DAPO,领先超过10分。VAPO训练过程表现出高度稳定与高效,在仅5000步内即达到最优性能,且多次独立运行均未出现训练崩溃。本研究聚焦于长思维链(long-CoT)推理,识别出价值基方法面临的三大问题:价值模型偏差、序列长度异质性以及奖励信号稀疏性。通过系统性设计,VAPO提供一体化解决方案,有效缓解上述挑战,显著提升长思维链推理任务表现。
原文摘要 · Abstract (English)
We present VAPO, Value-based Augmented Proximal Policy Optimization framework for reasoning models., a novel framework tailored for reasoning models within the value-based paradigm. Benchmarked the AIME 2024 dataset, VAPO, built on the Qwen 32B pre-trained model, attains a state-of-the-art score of $\mathbf{60.4}$. In direct comparison under identical experimental settings, VAPO outperforms the previously reported results of DeepSeek-R1-Zero-Qwen-32B and DAPO by more than 10 points. The training process of VAPO stands out for its stability and efficiency. It reaches state-of-the-art performance within a mere 5,000 steps. Moreover, across multiple independent runs, no training crashes occur, underscoring its reliability. This research delves into long chain-of-thought (long-CoT) reasoning using a value-based reinforcement learning framework. We pinpoint three key challenges that plague value-based methods: value model bias, the presence of heterogeneous sequence lengths, and the sparsity of reward signals. Through systematic design, VAPO offers an integrated solution that effectively alleviates these challenges, enabling enhanced performance in long-CoT reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。