arXiv:2509.25300cs.LGcs.AI2025-09ACL被引 18

研究大模型强化学习训练的扩展规律,发现模型越大越高效,但效率有饱和点。

Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning

  • 通过全系列Qwen2.5模型(0.5B到72B)实验,分析规模、数据量与算力的交互关系
  • 发现测试损失与算力、数据量呈稳健幂律关系,且大模型学习效率更高
  • 数据有限时重复使用高质量数据更有效,性能取决于优化步数而非样本多样性

尽管大语言模型(LLMs)预训练阶段的扩展规律已被广泛研究,但其在强化学习(RL)后训练中的表现仍不明确。本文系统性地实证研究了基于RL的后训练扩展行为,聚焦于数学推理能力。基于对全系列Qwen2.5密集模型(0.5B至72B)的实验,我们刻画了模型规模、数据量与计算预算如何共同影响性能。研究得出四个关键发现:1. 更大模型在算力和数据效率上均表现出更强的学习优势;2. 测试损失与算力、数据量之间的关系可由一个对基线与指令微调模型均稳健的预测幂律模型描述;3. 尽管大模型学习效率更高,但幂律中的学习效率项k(N)显示其随模型规模增长呈现潜在饱和趋势;4. 在数据受限情况下,重复使用高质量数据极为有效,最终性能主要由总优化步数决定,而非样本独特性。这些结果为通过强化学习高效扩展大模型推理能力提供了理论基础与实践指导。

原文摘要 · Abstract (English)

While scaling laws for large language models (LLMs) during pre-training have been extensively studied, their behavior under reinforcement learning (RL) post-training remains largely unexplored. This paper presents a systematic empirical investigation of scaling behaviors in RL-based post-training, with a particular focus on mathematical reasoning. Based on a set of experiments across the full Qwen2.5 dense model series (0.5B to 72B), we characterize how model scale, data volume, and computational budget interact to shape performance. Our analysis leads to four key findings: 1. Larger models consistently exhibit superior learning efficiency on both compute and data metrics. 2. The relationship between test loss, compute, and data can be modeled by a predictive power-law which is robust across both base and instruction-tuned models. 3. Although larger models exhibit higher learning efficiency, the analytical learning efficiency term k(N) in the power-law reveals a latent saturation trend in learning efficiency as model size continues to increase. 4. In data-constrained regimes, repeated reuse of high-quality data proves highly effective, as final performance is primarily governed by the total number of optimization steps rather than the uniqueness of samples. Collectively, these results provide a principled foundation and practical guidelines for efficiently scaling the reasoning capabilities of LLMs through RL post-training.

强化学习模型扩展数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。