arXiv:2502.11152math.OCcs.LG2025-02中稿 · publication in Mat…被引 1

分析深度线性网络正则化损失的误差界,揭示梯度与临界点距离的关系。

Error Bound Analysis for the Regularized Loss of Deep Linear Neural Networks

  • 通过闭式表达刻画临界点集合,建立正则化损失的误差界。
  • 在适度宽网络和正则化参数下,梯度范数可量化到临界点的距离。
  • 理论支持梯度下降线性收敛,适合研究优化机制的研究者阅读。

深度线性网络的优化基础近年受到广泛关注。然而,由于其固有的非凸性和层级结构,分析深度线性网络的损失函数仍具挑战性。本文研究了深度线性网络正则化平方损失在每个临界点附近的局部几何性质。具体而言,我们基于已有结果,获得了临界点集的闭式刻画,并在对网络宽度和正则化参数施加温和条件的情况下,建立了正则化损失的误差界。值得注意的是,该误差界将某点到临界点集的距离,以当前梯度范数的形式进行量化,可用于推导一阶方法的线性收敛性。为支持理论结果,我们进行了数值实验,表明当优化深度线性网络的正则化损失时,梯度下降能线性收敛至临界点。

原文摘要 · Abstract (English)

The optimization foundations of deep linear networks have recently received significant attention. However, due to their inherent non-convexity and hierarchical structure, analyzing the loss functions of deep linear networks remains a challenging task. In this work, we study the local geometry of the regularized squared loss of deep linear networks around each critical point. Specifically, we obtain a closed-form characterization of the critical point set building on existing results and establish an error bound for the regularized loss under mild conditions on network width and regularization parameters. Notably, this error bound quantifies the distance from a point to the critical point set in terms of the current gradient norm, which can be used to derive linear convergence of first-order methods. To support our theoretical findings, we conduct numerical experiments and demonstrate that gradient descent converges linearly to a critical point when optimizing the regularized loss of deep linear networks.

深度线性网络优化理论误差界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。