106B参数专家混合模型,强化学习训练后在多个领域超越更大模型。
INTELLECT-3: Technical Report
- 基于大规模强化学习,采用异步分布式框架训练12B活跃参数的专家混合模型。
- 在数学、代码、科学和推理任务上表现领先,优于许多更大模型。
- 开源完整训练基础设施,适合研究强化学习与智能体系统的开发者。
我们提出 INTELLECT-3,一个 1060 亿参数的专家混合模型(激活参数 120 亿),在端到端强化学习基础设施上进行大规模强化学习训练。该模型在数学、代码、科学和推理基准测试中达到与其规模相当的最先进水平,性能超越多个更大规模的前沿模型。我们开源了该模型及其完整的训练基础设施,包括强化学习框架、完整训练流程和由验证器库构建的丰富环境集合,来自我们的 Environments Hub 社区平台。为支持本工作,我们推出了 prime-rl——一个面向大规模异步强化学习的开源框架,可从单节点无缝扩展至数千张 GPU,专为多轮交互和工具使用等智能体强化学习设计。利用此栈,我们在 GLM-4.5-Air-Base 模型基础上运行 SFT 和强化学习训练,将强化学习训练扩展至 512 张 H200 显卡,保持高训练效率。
原文摘要 · Abstract (English)
We present INTELLECT-3, a 106B-parameter Mixture-of-Experts model (12B active) trained with large-scale reinforcement learning on our end-to-end RL infrastructure stack. INTELLECT-3 achieves state of the art performance for its size across math, code, science and reasoning benchmarks, outperforming many larger frontier models. We open-source the model together with the full infrastructure stack used to create it, including RL frameworks, complete recipe, and a wide collection of environments, built with the verifiers library, for training and evaluation from our Environments Hub community platform. Built for this effort, we introduce prime-rl, an open framework for large-scale asynchronous reinforcement learning, which scales seamlessly from a single node to thousands of GPUs, and is tailored for agentic RL with first-class support for multi-turn interactions and tool use. Using this stack, we run both SFT and RL training on top of the GLM-4.5-Air-Base model, scaling RL training up to 512 H200s with high training efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。