千亿参数模型实现突破性推理能力,开源释放顶尖思维智能。
Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model
- 通过分块动态拆分与令牌级修正,解决训练推理不一致问题。
- 在多项基准测试中达93.4至55.94分,获IMO2025银牌级表现。
- 首次开放万亿参数稀疏模型,推动大模型推理能力普惠化。
我们提出Ring-1T,首个开源的万亿级参数思维模型,拥有1万亿总参数,每令牌激活约500亿参数。训练此类模型面临前所未有的挑战,包括训练-推理不一致、回放处理效率低下及强化学习系统瓶颈。为此,我们提出三项相互关联的创新:(1) IcePop通过令牌级差异掩码与截断稳定RL训练,解决训练-推理失配导致的不稳定性;(2) C3PO++在令牌预算下动态划分长回放序列,显著提升时间效率;(3) ASystem是高性能强化学习框架,有效克服系统性瓶颈。Ring-1T在关键评测中表现卓越:AIME-2025得93.4分,HMMT-2025得86.72分,CodeForces得2088分,ARC-AGI-1得55.94分,更在IMO-2025达到银牌水平,彰显其强大推理能力。通过向社区开源完整1T参数MoE模型,我们为研究者提供直接访问前沿推理能力的通道。此项工作标志着大模型推理智能民主化的里程碑,确立了开源模型性能新基准。
原文摘要 · Abstract (English)
We present Ring-1T, the first open-source, state-of-the-art thinking model with a trillion-scale parameter. It features 1 trillion total parameters and activates approximately 50 billion per token. Training such models at a trillion-parameter scale introduces unprecedented challenges, including train-inference misalignment, inefficiencies in rollout processing, and bottlenecks in the RL system. To address these, we pioneer three interconnected innovations: (1) IcePop stabilizes RL training via token-level discrepancy masking and clipping, resolving instability from training-inference mismatches; (2) C3PO++ improves resource utilization for long rollouts under a token budget by dynamically partitioning them, thereby obtaining high time efficiency; and (3) ASystem, a high-performance RL framework designed to overcome the systemic bottlenecks that impede trillion-parameter model training. Ring-1T delivers breakthrough results across critical benchmarks: 93.4 on AIME-2025, 86.72 on HMMT-2025, 2088 on CodeForces, and 55.94 on ARC-AGI-1. Notably, it attains a silver medal-level result on the IMO-2025, underscoring its exceptional reasoning capabilities. By releasing the complete 1T parameter MoE model to the community, we provide the research community with direct access to cutting-edge reasoning capabilities. This contribution marks a significant milestone in democratizing large-scale reasoning intelligence and establishes a new baseline for open-source model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。