揭示深度神经网络训练中相变现象的几何根源
Phase transitions reveal hierarchical structure in deep neural networks
- 用鞍点解释深度网络训练中的相变现象
- 提出基于L2正则的快速算法,验证了模型连接性
- 发现不同数字类别的模型间存在可连通的几何路径
深度神经网络的训练依赖于模型在高维非凸损失曲面上收敛到优质极小值。然而,训练过程中的诸多现象仍不清晰。本文聚焦三个看似无关的现象:类统计物理的相变、鞍点普遍性以及模型连接性(用于模型合并)。我们统一将其归因于损失与误差景观的几何结构。理论上证明,深度网络学习中的相变由损失景观中的鞍点主导。基于此,提出一种简单、快速、易实现的算法,利用L2正则化探测误差景观。应用于MNIST数据集训练的DNN,高效找到了连接全局极小值的路径,验证了模型连接性。数值实验进一步表明,鞍点驱动了编码不同数字类别的模型间的转换。研究揭示了深度网络关键训练现象的几何起源,并发现类似统计物理相变的精度基底层级结构。
原文摘要 · Abstract (English)
Training Deep Neural Networks relies on the model converging on a high-dimensional, non-convex loss landscape toward a good minimum. Yet, much of the phenomenology of training remains ill understood. We focus on three seemingly disparate observations: the occurrence of phase transitions reminiscent of statistical physics, the ubiquity of saddle points, and phenomenon of mode connectivity relevant for model merging. We unify these within a single explanatory framework, the geometry of the loss and error landscapes. We analytically show that phase transitions in DNN learning are governed by saddle points in the loss landscape. Building on this insight, we introduce a simple, fast, and easy to implement algorithm that uses the L2 regularizer as a tool to probe the geometry of error landscapes. We apply it to confirm mode connectivity in DNNs trained on the MNIST dataset by efficiently finding paths that connect global minima. We then show numerically that saddle points induce transitions between models that encode distinct digit classes. Our work establishes the geometric origin of key training phenomena in DNNs and reveals a hierarchy of accuracy basins analogous to phases in statistical physics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。