提出随机高斯-牛顿法在过参数化模型中的非渐近优化与泛化理论
Non-Asymptotic Optimization and Generalization Bounds for Stochastic Gauss-Newton in Overparameterized Models
- 基于变度量分析建立有限时间收敛界,显式关联批大小、网络宽度与深度
- 在过参数化下给出非渐近泛化界,揭示曲率、批大小和过参数化的影响
- 发现优化路径上高斯-牛顿矩阵最小特征值越大,泛化越优,适合研究泛化机制
深度学习中一个核心问题是高阶优化方法如何影响泛化性能。本文针对带有Levenberg-Marquardt阻尼和小批量采样的随机高斯-牛顿(SGN)方法,在回归设置下分析了具有光滑激活函数的过参数化深度神经网络的训练过程。理论贡献有两方面:第一,通过参数空间中的变度量分析,建立了有限时间收敛界,并显式刻画了批大小、网络宽度和深度的影响;第二,利用过参数化下的均匀稳定性,推导出SGN的非渐近泛化界,定量分析了曲率、批大小和过参数化对泛化性能的作用。理论结果识别出一个有利的泛化区域:在优化路径上,若高斯-牛顿矩阵的最小特征值较大,则稳定性界更紧。
原文摘要 · Abstract (English)
An important question in deep learning is how higher-order optimization methods affect generalization. In this work, we analyze a stochastic Gauss-Newton (SGN) method with Levenberg-Marquardt damping and mini-batch sampling for training overparameterized deep neural networks with smooth activations in a regression setting. Our theoretical contributions are twofold. First, we establish finite-time convergence bounds via a variable-metric analysis in parameter space, with explicit dependencies on the batch size, network width and depth. Second, we derive non-asymptotic generalization bounds for SGN using uniform stability in the overparameterized regime, characterizing the impact of curvature, batch size, and overparameterization on generalization performance. Our theoretical results identify a favorable generalization regime for SGN in which a larger minimum eigenvalue of the Gauss-Newton matrix along the optimization path yields tighter stability bounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。