用离线强化学习提升大模型数学推理能力,效果超越在线方法。
PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning
- 基于离线强化学习训练320亿参数模型,提升稳定性与效率。
- 在AIME 2024和2025测试集上分别达到90.9%和85.6%准确率。
- 适合关注大模型推理能力优化的研究者与开发者。
我们提出PCL-Reasoner-V1.5,一个320亿参数的大语言模型,专用于数学推理。该模型基于Qwen2.5-32B构建,通过监督微调(SFT)后接强化学习(RL)进行优化。核心创新在于提出的离线强化学习方法,相比标准在线RL方法(如GRPO)具有更优的训练稳定性和效率。模型在基于Qwen2.5-32B的后训练模型中表现领先,在AIME 2024和AIME 2025测试集上平均准确率分别达到90.9%和85.6%。所有实验均在华为Ascend 910C NPUs上完成。本工作验证了离线强化学习作为大模型推理能力提升的稳定高效范式。
原文摘要 · Abstract (English)
We present PCL-Reasoner-V1.5, a 32-billion-parameter large language model (LLM) for mathematical reasoning. The model is built upon Qwen2.5-32B and refined via supervised fine-tuning (SFT) followed by reinforcement learning (RL). A central innovation is our proposed offline RL method, which provides superior training stability and efficiency over standard online RL methods such as GRPO. Our model achieves state-of-the-art performance among models post-trained on Qwen2.5-32B, attaining average accuracies of 90.9% on AIME 2024 and 85.6% on AIME 2025. Our work demonstrates offline RL as a stable and efficient paradigm for advancing reasoning in LLMs. All experiments were conducted on Huawei Ascend 910C NPUs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。