arXiv:2505.24452cs.LG2025-05NeurIPS被引 1

提出统一学习率调度方法,高效适配不同模型与预算。

Stepsize anything: A unified learning rate schedule for budgeted-iteration training

  • 基于预算感知优化框架,推导出单参数控制的UBA调度
  • 在多种模型和预算下均超越常用调度方案,提升训练效率
  • 理论证明收敛性并给出参数选择指南,适合资源受限场景

计算成本上升与资源有限性凸显了预算迭代训练的重要性,即在预设迭代次数内实现最优学习。尽管学习率调度对网络性能有决定性影响,尤其在预算约束场景中,其设计仍依赖经验,缺乏理论支撑,且需大量试错,导致训练低效。本文提出统一预算感知(UBA)调度,一种理论上严谨的学习率调度方法,在不同架构与任务下,于各类训练预算中持续优于常见调度方案。首先,我们构建新颖的预算感知优化框架,显式考虑对景观曲率变化的鲁棒性;由此推导出由单一超参数φ控制的UBA调度,实现灵活性与简洁性的平衡,无需针对每类网络进行数值优化。此外,我们建立φ与条件数间的理论关联,增强方法可解释性。进一步证明了不同φ值下的收敛性,并通过理论分析与实证结果提供实用参数选择建议。大量实验表明,无论视觉或语言任务、不同网络架构(如ResNet、OLMo)与规模,在多种训练迭代预算下,UBA均一致超越现有常用调度方案。

原文摘要 · Abstract (English)

The expanding computational costs and limited resources underscore the critical need for budgeted-iteration training, which aims to achieve optimal learning within predetermined iteration budgets. While learning rate schedules fundamentally govern the performance of different networks and tasks, particularly in budgeted-iteration scenarios, their design remains largely heuristic, lacking theoretical foundations. In addition, the optimal learning rate schedule requires extensive trial-and-error selection, making the training process inefficient. In this work, we propose the Unified Budget-Aware (UBA) schedule, a theoretically grounded learning rate schedule that consistently outperforms commonly-used schedules among diverse architectures and tasks under different constrained training budgets. First, we bridge the gap by constructing a novel training budget-aware optimization framework, which explicitly accounts for the robustness to landscape curvature variations. From this framework, we derive the UBA schedule, controlled by a single hyper-parameter φthat provides a trade-off between flexibility and simplicity, eliminating the need for per-network numerical optimization. Moreover, we establish a theoretical connection between φand the condition number, adding interpretation and justification to our approach. Besides, we prove the convergence for different values of φ. We offer practical guidelines for its selection via theoretical analysis and empirical results. Extensive experimental results show that UBA consistently surpasses the commonly-used schedules across diverse vision and language tasks, spanning network architectures (e.g., ResNet, OLMo) and scales, under different training-iteration budgets.

学习率调度预算训练优化理论深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。