arXiv:2608.09523cs.LG2026-08

用对偶理论重新定义神经网络优化,解释为何随机梯度下降有效

Generalized Convexity and Smoothness via Conjugate Duality: Optimization Theory for Deep Neural Networks

论文配图:Generalized Convexity and Smoothness via Conjugate Duality: Optimization Theory for Deep Neural Networks
图 1 · 摘自论文原文
  • 基于共轭对偶提出广义凸性与光滑性,统一处理非凸非光滑问题
  • 证明广义梯度下降最优学习率恰为1,给出基于梯度能量的收敛速率
  • 揭示网络结构、批量大小等如何通过梯度相关因子影响训练收敛

深度神经网络(DNN)训练采用随机梯度下降(SGD)及其变体取得了出色的实证性能,但经典优化理论难以完全解释这一成功。其原因在于传统分析依赖于可微性、凸性或光滑性等假设,而这些往往不适用于DNN目标函数。本文通过勒让德函数与凸共轭,将经典凸性与光滑性推广,提出$\ ext{H}(ψ)$-凸性和$\ ext{H}(Ψ)$-光滑性,统一处理凸与非凸、光滑与非光滑目标,并揭示广义光滑性与凸性间的自然对偶关系。基于此,引入广义梯度下降(GD)与广义SGD,理论上证明广义GD的最优学习率为1,并推导出两种优化器的梯度能量收敛速率。进一步将DNN训练重构为复合优化问题,表明收敛依赖于同时降低梯度能量并控制网络雅可比矩阵的诱导范数。为刻画网络架构与训练配置的实际影响,引入梯度相关因子与模型容量风险,定量分析了架构设计、批量大小及模型容量对训练收敛的作用。在多种网络结构、数据集、优化器与损失函数上的大量实验验证了理论边界,且理论预测与实际训练动态高度一致。

原文摘要 · Abstract (English)

Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success. This limitation arises because conventional analyses rely on assumptions such as differentiability, convexity, or smoothness, which are often violated by DNN objectives. In this paper, we establish a unified optimization framework for DNN training by generalizing classical convexity and smoothness through Legendre functions and convex conjugation. Specifically, we introduce $\mathcal{H}(ψ)$-convexity and $\mathcal{H}(Ψ)$-smoothness, which unify convex and non-convex as well as smooth and non-smooth objectives within a single formalism and reveal a natural duality between generalized smoothness and convexity. Building on these generalized properties, we introduce generalized gradient descent (GD) and generalized SGD through convex conjugation. We theoretically prove that generalized GD admits an optimal learning rate of exactly $1$, and derive rigorous gradient-energy-based convergence rates for both proposed optimizers. We further reformulate DNN training as a composite optimization problem, demonstrating that its convergence relies on jointly reducing the gradient energy and controlling the induced norm of the network Jacobian. To characterize the practical influences of network architectures and training configurations, we introduce the gradient correlation factor and model capacity risk, and quantitatively analyze how architectural designs, batch size, and model capacity shape training convergence. Extensive experiments across diverse network architectures, datasets, optimizers, and loss functions validate our theoretical bounds and demonstrate precise alignment between our theoretical predictions and empirical training dynamics.

优化理论神经网络梯度下降共轭对偶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。