30B参数模型在数学编程竞赛中达顶尖水平,仅用20分之一参数量
Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
- 分阶段强化学习+多领域在线蒸馏,提升推理与智能体能力
- 数学与编程竞赛表现逼近顶级开源模型,参数量仅其1/20
- 开源完整模型与训练数据,适合研究高效大模型设计
我们提出Nemotron-Cascade 2,一个开源的300亿参数稀疏激活模型(激活参数30亿),在推理与智能体能力上达到业界领先水平。尽管规模紧凑,其数学与编程推理表现接近前沿开源模型。它是继DeepSeekV3.2-Speciale-671B-A37B之后,第二个在2025年国际数学奥林匹克(IMO)、国际信息学奥林匹克(IOI)及ICPC世界总决赛中获得金牌等级表现的开源大模型,展现出20倍更高的智能密度。相比Nemotron-Cascade 1,关键技术进步包括:在精心构建的数据集上完成SFT后,大幅扩展分阶段强化学习(Cascade RL)覆盖范围至更广泛的推理与智能体任务;同时引入多领域在线蒸馏机制,从各领域的最强中间教师模型持续蒸馏,有效恢复基准性能下降并保持持续优化。我们公开发布模型检查点与训练数据集。
原文摘要 · Abstract (English)
We introduce Nemotron-Cascade 2, an open 30B MoE model with 3B activated parameters that delivers best-in-class reasoning and strong agentic capabilities. Despite its compact size, its mathematical and coding reasoning performance approaches that of frontier open models. It is the second open-weight LLM, after DeepSeekV3.2-Speciale-671B-A37B, to achieve Gold Medal-level performance in the 2025 International Mathematical Olympiad (IMO), the International Olympiad in Informatics (IOI), and the ICPC World Finals, demonstrating remarkably high intelligence density with 20x fewer parameters. In contrast to Nemotron-Cascade 1, the key technical advancements are as follows. After SFT on a meticulously curated dataset, we substantially expand Cascade RL to cover a much broader spectrum of reasoning and agentic domains. Furthermore, we introduce multi-domain on-policy distillation from the strongest intermediate teacher models for each domain throughout the Cascade RL process, allowing us to efficiently recover benchmark regressions and sustain strong performance gains along the way. We release the collection of model checkpoint and training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。