arXiv:2602.07145cs.LGcs.CL2026-02中稿 · ICLR被引 4

发现深度学习训练初期即呈现弱凸性,可据此推导学习率缩放规律。

Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate

  • 基于凸性分析,提出损失动态的上界预测方法。
  • 在80倍训练时长和70倍模型规模下仍能准确外推学习率与损失。
  • 为优化器调参提供理论依据,适合模型训练调优研究者。

深度学习具有非凸损失函数,其优化过程难以分析或控制。然而,实证表明其动态行为在各类任务、模型、优化器及超参数设置下均表现出类似凸性的特征。本文考察凸性和Lipschitz连续性在深度学习中的适用性,旨在通过学习率调度精确控制损失动态。结果表明,深度学习在训练初期很快进入弱凸状态,损失可通过最后一轮迭代的上界进行预测,进而指导最优学习率的缩放。基于凸性视角,建立了学习率与损失的缩放规律,可在训练周期延长80倍、模型规模扩大70倍的情况下实现有效外推。

原文摘要 · Abstract (English)

Deep learning has non-convex loss landscape and its optimization dynamics is hard to analyze or control. Nevertheless, the dynamics can be empirically convex-like across various tasks, models, optimizers, hyperparameters, etc. In this work, we examine the applicability of convexity and Lipschitz continuity in deep learning, in order to precisely control the loss dynamics via the learning rate schedules. We illustrate that deep learning quickly becomes weakly convex after a short period of training, and the loss is predicable by an upper bound on the last iterate, which further informs the scaling of optimal learning rate. Through the lens of convexity, we build scaling laws of learning rates and losses that extrapolate as much as 80X across training horizons and 70X across model sizes.

深度学习优化理论学习率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。