arXiv:2603.15958cs.LG2026-03被引 7

从优化理论出发,推导出学习率与批量大小的最优缩放规律。

Deriving Hyperparameter Scaling Laws via Modern Optimization Theory

  • 基于线性最小化预言机框架,分析现代一阶优化器收敛性。
  • 推导出学习率、动量和批量大小随训练预算的幂律缩放公式。
  • 揭示动量与批量大小的协同作用,为高效训练提供新思路。

超参数迁移已成为大规模训练的重要组成部分。现有方法如muP主要关注模型规模间的迁移,而批量大小和训练时长的迁移通常依赖于基于时间尺度保持、二次代理和连续时间近似的经验缩放规则。本文通过近期基于线性最小化预言机(LMO)框架的收敛界,研究现代一阶优化器的超参数缩放规律,该框架包含归一化SGD、signSGD(近似Adam)和Muon。将文献中的界作为代理函数,在不同调参策略下最小化,得到学习率、动量和批量大小关于迭代次数或令牌预算的闭式幂律调度。在固定模型规模的前提下,本分析统一解释了文献中的多数观察结果,并指明未来研究方向。特别地,结果强调了动量与批量大小缩放之间的交互作用,暗示最优性能可能通过多种缩放策略实现。

原文摘要 · Abstract (English)

Hyperparameter transfer has become an important component of modern large-scale training recipes. Existing methods, such as muP, primarily focus on transfer between model sizes, with transfer across batch sizes and training horizons often relying on empirical scaling rules informed by insights from timescale preservation, quadratic proxies, and continuous-time approximations. We study hyperparameter scaling laws for modern first-order optimizers through the lens of recent convergence bounds for methods based on the Linear Minimization Oracle (LMO), a framework that includes normalized SGD, signSGD (approximating Adam), and Muon. Treating bounds in recent literature as a proxy and minimizing them across different tuning regimes yields closed-form power-law schedules for learning rate, momentum, and batch size as functions of the iteration or token budget. Our analysis, holding model size fixed, recovers most insights and observations from the literature under a unified and principled perspective, with clear directions open for future research. Our results draw particular attention to the interaction between momentum and batch-size scaling, suggesting that optimal performance may be achieved with several scaling strategies.

超参数优化理论缩放规律深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。