提出新理论分析框架,揭示了Lookahead-SGD在泛化上的优势。
Generalization and Optimization of SGD with Lookahead
- 基于平均模型稳定性,突破传统Lipschitz假设限制
- 证明凸问题下泛化误差随批量大小线性下降
- 为优化与泛化关系提供更真实、可解释的理论支持
Lookahead优化器通过双权重更新机制提升深度学习模型性能,但其泛化能力缺乏充分理论支撑。现有分析常依赖全局Lipschitz连续性等强假设,且未能完整刻画优化与泛化的关系。本文针对小批量SGD下的Lookahead优化器,开展严格的稳定性和泛化性分析。利用平均模型稳定性,推导出无需全局Lipschitz假设的凸与强凸问题的泛化界。结果表明,在凸设置下,泛化误差随批量大小呈线性加速下降。
原文摘要 · Abstract (English)
The Lookahead optimizer enhances deep learning models by employing a dual-weight update mechanism, which has been shown to improve the performance of underlying optimizers such as SGD. However, most theoretical studies focus on its convergence on training data, leaving its generalization capabilities less understood. Existing generalization analyses are often limited by restrictive assumptions, such as requiring the loss function to be globally Lipschitz continuous, and their bounds do not fully capture the relationship between optimization and generalization. In this paper, we address these issues by conducting a rigorous stability and generalization analysis of the Lookahead optimizer with minibatch SGD. We leverage on-average model stability to derive generalization bounds for both convex and strongly convex problems without the restrictive Lipschitzness assumption. Our analysis demonstrates a linear speedup with respect to the batch size in the convex setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。