arXiv:2412.14031math.OCcs.AI2024-12被引 2

提出用黎曼优化视角分析高斯-牛顿法,实现神经网络训练的加速收敛。

A Riemannian Optimization Perspective of the Gauss-Newton Method for Feedforward Neural Networks

  • 将参数空间的高斯-牛顿流映射到函数空间的黎曼流,构建几何优化框架。
  • 在欠参和过参情形下均实现与条件数无关的几何收敛率。
  • 适合研究优化算法性能、设计高效训练方法的学者参考。

本文为具有光滑激活函数的前馈神经网络训练中高斯-牛顿法建立了非渐近收敛界。在欠参情形下,参数空间中的高斯-牛顿梯度流在函数空间的低维嵌入子流形上诱导出黎曼梯度流。借助黎曼优化工具,在适当输出缩放下,建立了损失函数的测地线Polyak-Lojasiewicz条件与Lipschitz光滑性,实现了对类内最优预测器的显式几何收敛,且收敛速率独立于格拉姆矩阵的条件数。在过参情形下,提出自适应、曲率感知的正则化调度策略,确保以独立于神经切向核最小特征值的速率快速几何收敛至全局最优,并在局部独立于损失的强凸模量。这些结果表明,当一阶方法因核矩阵病态或损失景观复杂而收敛缓慢时,高斯-牛顿法仍能实现加速收敛。

原文摘要 · Abstract (English)

In this work, we establish non-asymptotic convergence bounds for the Gauss-Newton method in training neural networks with smooth activations. In the underparameterized regime, the Gauss-Newton gradient flow in parameter space induces a Riemannian gradient flow on a low-dimensional embedded submanifold of the function space. Using tools from Riemannian optimization, we establish geodesic Polyak-Lojasiewicz and Lipschitz-smoothness conditions for the loss under appropriately chosen output scaling, yielding geometric convergence to the optimal in-class predictor at an explicit rate independent of the conditioning of the Gram matrix. In the overparameterized regime, we propose adaptive, curvature-aware regularization schedules that ensure fast geometric convergence to a global optimum at a rate independent of the minimum eigenvalue of the neural tangent kernel and, locally, of the modulus of strong convexity of the loss. These results demonstrate that Gauss-Newton achieves accelerated convergence rates in settings where first-order methods exhibit slow convergence due to ill-conditioned kernel matrices and loss landscapes.

优化算法黎曼优化神经网络收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。