arXiv:2603.21191cs.LGmath.OC2026-03被引 6

揭示批量大小对随机条件梯度方法的影响机制,指导大模型训练的批处理策略。

On the Role of Batch Size in Stochastic Conditional Gradient Methods

  • 基于μ-KL条件分析批量大小与步长、噪声的交互关系。
  • 批量过大时性能会饱和甚至下降,存在最优批量阈值。
  • 提出动态增大批量和序列长度的自适应策略,适合大模型训练。

我们在μ-柯尔达-洛贾斯耶维奇(μ-KL)条件下研究了批量大小在随机条件梯度方法中的作用。针对基于动量的随机条件梯度算法(如Scion),我们提出了新分析,明确捕捉了步长、批量大小与随机噪声之间的相互作用。研究发现:增加批量大小初期可提升优化精度,但超过临界阈值后,性能改善趋于饱和,甚至在固定令牌预算下出现退化。理论预测了最优步长的量级,并与大规模训练中的经验实践高度一致。基于这些洞察,我们推导出批量大小与步长选择的合理准则,并提出一种训练过程中逐步增加批量大小和序列长度的自适应策略,同时保证收敛性。在NanoGPT上的实验验证了理论预测,展示了预期的缩放规律。总体而言,本研究为理解随机条件梯度方法中批量大小的缩放行为提供了理论框架,并为大规模优化中的高效训练调度提供指导。

原文摘要 · Abstract (English)

We study the role of batch size in stochastic conditional gradient methods under a $μ$-Kurdyka-Łojasiewicz ($μ$-KL) condition. Focusing on momentum-based stochastic conditional gradient algorithms (e.g., Scion), we derive a new analysis that explicitly captures the interaction between stepsize, batch size, and stochastic noise. Our study reveals a regime-dependent behavior: increasing the batch size initially improves optimization accuracy but, beyond a critical threshold, the benefits saturate and can eventually degrade performance under a fixed token budget. Notably, the theory predicts the magnitude of the optimal stepsize and aligns well with empirical practices observed in large-scale training. Leveraging these insights, we derive principled guidelines for selecting the batch size and stepsize, and propose an adaptive strategy that increases batch size and sequence length during training while preserving convergence guarantees. Experiments on NanoGPT are consistent with the theoretical predictions and illustrate the emergence of the predicted scaling regimes. Overall, our results provide a theoretical framework for understanding batch size scaling in stochastic conditional gradient methods and offer guidance for designing efficient training schedules in large-scale optimization.

优化算法批量大小条件梯度大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。