arXiv:2605.24316cs.LG2026-05被引 1

揭示动态小批量SGD在降维线性回归中的规模定律与误差机制

Scaling Laws for Dynamic Mini-Batch SGD in Sketched Linear Regression

论文配图:Scaling Laws for Dynamic Mini-Batch SGD in Sketched Linear Regression
图 1 · 摘自论文原文
  • 基于谱幂律和源条件,分析单遍与多遍小批量优化过程
  • 发现两种误差:确定性误差依赖全程轨迹,随机误差仅保留末段记忆
  • 提出最优迭代分配规则,适用于模型训练资源规划

小批量训练是大规模优化的核心,但其在统计规模定律中的作用仍不明确。本文研究在谱幂律和源条件下的降维线性回归中的一遍与多遍小批量随机梯度下降。分析揭示了由预热-稳定-衰减调度引发的双时域现象:确定性学习受完整优化轨迹控制,而随机误差仅保留较短的末端记忆。对于动态批次调度,各批次大小通过影响加权汇总量体现其对最终风险的影响。因此,在固定更新时域下,分批不影响近似与优化偏差规律,但调控一遍方差及多遍中相对于全批量梯度下降的波动。我们得到了匹配的一遍方差界和几乎匹配的多遍波动界,恢复了静态批次与全批量行为作为特例,并推导出在固定迭代预算下的最优平方根分配规则。这些结果揭示了WSD时域分离与最终风险影响是动态小批量扩展的核心机制。

原文摘要 · Abstract (English)

Mini-batching is central to large-scale optimization, yet its role in statistical scaling laws remains limited. We study one-pass and multi-pass batch SGD for sketched linear regression under power-law spectral and source conditions. Our analysis reveals a two-horizon phenomenon induced by warmup--stable--decay schedules: deterministic learning is governed by the full optimization trajectory, while stochastic error retains only a shorter terminal memory. For dynamic batch schedules, the individual batch sizes enter through influence-weighted summaries that measure how strongly each update affects the final risk. Consequently, batching leaves the approximation and optimization-bias laws unchanged at a fixed update horizon, but controls the one-pass variance and the multi-pass fluctuation around full-batch gradient descent. We obtain matching one-pass variance bounds and nearly matching multi-pass fluctuation bounds, recover static-batch and full-batch behavior as special cases, and derive an oracle square-root rule for allocating a fixed iteration budget. These results identify WSD horizon separation and final-risk influence as the mechanisms governing dynamic mini-batch scaling.

优化算法统计学习规模定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。