训练1万亿参数模型实现自主推理,揭示规模化带来的新认知能力。
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

- 通过算法与系统优化,实现1万亿参数的零样本强化学习训练。
- 模型在7个数学基准上表现优异,且推理过程更结构化、简洁。
- 规模扩大后自发产生自验证、并行推理等高级认知行为,无需人工设计规则。
无人类标注数据的可验证奖励强化学习(即零样本强化学习)已成为激发链式思维推理的强大范式。然而受限于计算资源,现有研究多集中于小模型,大规模下的训练动态与涌现能力尚未探索。为突破此限制,我们致力于从大模型中诱导高质量推理行为。发现简单扩展常导致可读性差、重复冗余及推理深度不足。为此,提出稳定高效的训练流水线,融合截断重要性采样、训练-推理比例修正和混合精度控制等优化策略。实验揭示三个关键发现:(1)扩展至1万亿参数显著提升样本效率与性能上限;(2)训练过程经历初始发现阶段与后续精炼阶段;(3)模型自发形成类人特征、结构化格式、自我验证、并行推理与情境焦虑等高级认知行为,使手工设计启发式方法变得多余。在七个数学基准上,Ring-2.5-1T-Zero表现竞争力。为进一步评估链式思维质量,提出涵盖可理解性、可复现性与效率的三维评估框架,模型在生成结构化、紧凑推理路径方面优势明显。通过共享观测到的涌现现象,期望为社区提供1万亿规模下规模化行为的深层洞察。
原文摘要 · Abstract (English)
Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small models, leaving the training dynamics and emergent capabilities at a large scale unexplored. To meaningfully explore this frontier, we aim to elicit high-quality reasoning behaviors from the model. However, we find that naive scaling often suffers from poor readability, token redundancy, and a lack of adaptive reasoning depth. To address these challenges, we present a stable and efficient training pipeline, incorporating algorithmic and system optimizations such as clipped importance sampling, training-inference ratio correction, and mixed-precision control. Our experiments offer three key findings that validate the "bitter lesson" of scaling: (1) scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; (2) the training process progresses sequentially through an initial discovery phase followed by a sharpening phase; and (3) the model spontaneously develops advanced cognitive behaviors, including anthropomorphism, structured formatting, self-verification, parallel reasoning, and context anxiety, rendering hand-crafted heuristics redundant. Evaluated on seven mathematical benchmarks, Ring-2.5-1T-Zero achieves competitive performance. Additionally, to assess CoT quality beyond final-answer correctness, we propose a structured evaluation framework across three dimensions: comprehensibility, reproducibility, and efficiency, where our model demonstrates clear advantages in producing structured and concise reasoning traces. By sharing our observed emergent phenomena, we hope to provide the community with deeper insights into scaling behaviors, particularly at the 1-trillion scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。