arXiv:2411.02139cs.LGstat.ML2024-11NeurIPS被引 7

解析神经网络中高斯-牛顿矩阵的条件数,揭示其优化特性。

Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks

  • 从理论推导深度线性网络的高斯-牛顿矩阵条件数上界
  • 证明条件数随网络深度与宽度增长而加剧
  • 适用于研究优化器设计与网络结构影响的科研人员

高斯-牛顿(GN)矩阵在机器学习中具有重要作用,尤其作为多种自适应优化方法的预条件矩阵以加速优化过程,同时可提供对神经网络优化景观的关键洞察。在深度神经网络中,理解GN矩阵需分析不同权重矩阵间的交互以及数据引入的依赖关系,使得其分析极具挑战性。本文首次从理论上刻画神经网络中GN矩阵的条件性,建立了任意深度与宽度的深度线性网络中GN条件数的紧致上界,并拓展至两层ReLU网络。进一步分析了残差连接与卷积层等架构组件的影响。最后通过实验验证了这些上界,并揭示了所分析组件对条件数的实际影响。

原文摘要 · Abstract (English)

The Gauss-Newton (GN) matrix plays an important role in machine learning, most evident in its use as a preconditioning matrix for a wide family of popular adaptive methods to speed up optimization. Besides, it can also provide key insights into the optimization landscape of neural networks. In the context of deep neural networks, understanding the GN matrix involves studying the interaction between different weight matrices as well as the dependencies introduced by the data, thus rendering its analysis challenging. In this work, we take a first step towards theoretically characterizing the conditioning of the GN matrix in neural networks. We establish tight bounds on the condition number of the GN in deep linear networks of arbitrary depth and width, which we also extend to two-layer ReLU networks. We expand the analysis to further architectural components, such as residual connections and convolutional layers. Finally, we empirically validate the bounds and uncover valuable insights into the influence of the analyzed architectural components.

优化理论神经网络条件数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。