arXiv:2603.26954cs.LGmath.ST2026-03被引 1

分析双阶段优化器在高维线性回归中的表现,揭示其噪声与信号的平衡机制。

High dimensional theory of two-phase optimizers

  • 提出LA-DiLoCo模型,分阶段优化并同步参数
  • 单机版比SGD更优,多机版噪声可调参缓解
  • 动量叠加实现非线性谱变换,加速训练

随着训练规模扩大,部分异步双阶段优化器(在本地优化后跨工作节点同步)重新受到关注。近期研究发现,其中一种算法DiLoCo的单工作节点版本表现出色,可作为同步优化器使用。本文针对DiLoCo家族中的一个简单成员LA-DiLoCo,在高维线性回归问题上进行分析。结果表明,单工作节点版本LA在信号与噪声之间提供不同于SGD的权衡,适用于多种场景;多工作节点版本产生的噪声更多,但可通过合适超参数调节。进一步分析了带动量的SLA——LA加动量——发现叠加两个动量算子能通过非线性变换“有效”海森谱实现加速,且对Nesterov动量最优。总体表明,双阶段优化器为理解与改进训练算法提供了新范式。

原文摘要 · Abstract (English)

The trend towards larger training setups has brought a renewed interest in partially asynchronous two-phase optimizers which optimize locally and then synchronize across workers. Additionally, recent work suggests that the one-worker version of one of these algorithms, DiLoCo, shows promising results as a (synchronous) optimizer. Motivated by these studies we present an analysis of LA-DiLoCo, a simple member of the DiLoCo family, on a high-dimensional linear regression problem. We show that the one-worker variant, LA, provides a different tradeoff between signal and noise than SGD, which is beneficial in many scenarios. We also show that the multi-worker version generates more noise than the single worker version, but that this additional noise generation can be ameliorated by appropriate choice of hyperparameters. We conclude with an analysis of SLA -- LA with momentum -- and show that stacking two momentum operators gives an opportunity for acceleration via a non-linear transformation of the "effective'' Hessian spectrum, which is maximized for Nesterov momentum. Altogether our results show that two-phase optimizers represent a fruitful new paradigm for understanding and improving training algorithms.

优化器高维分析动量分布式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。