arXiv:2604.14669cs.LGmath.DS2026-04

揭示零阶优化在深度学习中的稳定边界与隐式正则化机制

Zeroth-Order Optimization at the Edge of Stability

论文配图:Zeroth-Order Optimization at the Edge of Stability
图 1 · 摘自论文原文
  • 基于两点估计器,推导出零阶优化的精确步长稳定性条件
  • 全批量零阶方法始终运行在稳定边界附近,验证理论预测
  • 零阶与一阶优化的正则化机制不同:前者主要约束海塞迹

当梯度不可用或计算成本过高时,零阶(ZO)方法广泛应用于黑箱学习和大模型内存高效微调,但其在深度学习中的优化动态仍缺乏深入研究。本文针对基于标准两点估计器的一类零阶方法,给出了精确的均方线性稳定性步长条件。分析表明,与仅依赖最大海塞特征值的一阶方法不同,零阶方法的均方稳定性取决于整个海塞谱。由于实际神经网络训练中无法计算完整海塞谱,我们进一步推导出仅依赖最大特征值和海塞迹的可计算稳定性上界。实验发现,全批量零阶方法(ZO-GD、ZO-GDM、ZO-Adam)始终在预测的稳定性边界附近稳定运行,涵盖多种深度学习任务。结果揭示了零阶方法特有的隐式正则化效应:大步长主要抑制海塞迹,而一阶方法则抑制最大特征值。

原文摘要 · Abstract (English)

Zeroth-order (ZO) methods are widely used when gradients are unavailable or prohibitively expensive, including black-box learning and memory-efficient fine-tuning of large models, yet their optimization dynamics in deep learning remain underexplored. In this work, we provide an explicit step size condition that exactly captures the (mean-square) linear stability of a family of ZO methods based on the standard two-point estimator. Our characterization reveals a sharp contrast with first-order (FO) methods: whereas FO stability is governed solely by the largest Hessian eigenvalue, mean-square stability of ZO methods depends on the entire Hessian spectrum. Since computing the full Hessian spectrum is infeasible in practical neural network training, we further derive tractable stability bounds that depend only on the largest eigenvalue and the Hessian trace. Empirically, we find that full-batch ZO methods operate at the edge of stability: ZO-GD, ZO-GDM, and ZO-Adam consistently stabilize near the predicted stability boundary across a range of deep learning training problems. Our results highlight an implicit regularization effect specific to ZO methods, where large step sizes primarily regularize the Hessian trace, whereas in FO methods they regularize the top eigenvalue.

零阶优化稳定性分析海塞谱正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。