用简单统一方法让模型达到奥数金牌水平,关键在自我检查与强化训练。
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling

- 通过逆困惑度课程微调,让模型学会严谨证明与自检。
- 经两阶段强化学习与测试时缩放,解决超长推理轨迹问题。
- 适合追求高阶数学物理推理的开发者与研究者使用。
近期推理模型在长程数学与科学问题求解上取得显著进展,已有系统达到国际数学奥林匹克(IMO)和国际物理奥林匹克(IPhO)金牌水平。本文提出一种简单且统一的方案,将后训练推理骨干转化为严谨的奥赛级求解器。该方案首先使用逆困惑度课程进行监督微调(SFT),以培养严谨的证明搜索与自检行为;随后通过两阶段强化学习(RL)流程,从可验证奖励逐步过渡到精细的证明级RL;最后利用测试时缩放进一步提升性能。基于此方法,我们对一个30B-A3B骨干模型在约34万条、每条小于8000词的轨迹上进行SFT,再进行200步强化学习。所获模型SU-01可在超过10万词的推理轨迹上稳定运行,并在数学与物理奥赛中达到金牌水平,涵盖IMO 2025/USAMO 2026及IPhO 2024/2025。此外,其科学推理能力在数学与物理之外领域也展现出强泛化性。
原文摘要 · Abstract (English)
Recent progress in reasoning models has substantially advanced long-horizon mathematical and scientific problem solving, with several systems now reaching gold-medal-level performance on International Mathematical Olympiad (IMO) and International Physics Olympiad (IPhO) problems. In this paper, we introduce a simple and unified recipe for converting a post-trained reasoning backbone into a rigorous olympiad-level solver. The recipe first uses a reverse-perplexity curriculum for SFT to instill rigorous proof-search and self-checking behaviors, then scales these behaviors through a two-stage RL pipeline that progresses from RL with verifiable rewards to more delicate proof-level RL, and finally boosts solving performance with test-time scaling. Applying this recipe, we train a 30B-A3B backbone with SFT on around 340K sub-8K-token trajectories followed by 200 RL steps. The resulting model, SU-01, supports stable reasoning on difficult problems with trajectories exceeding 100K tokens, while achieving gold-medal-level performance on mathematical and physical olympiad competitions, including IMO 2025/USAMO 2026 and IPhO 2024/2025. It also demonstrates strong generalization of scientific reasoning to domains beyond mathematics and physics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。