发现学习率影响神经元拓扑结构,小则保持复杂性,大则简化结构。
Topological Invariance and Breakdown in Learning
- 基于排列等变学习规则,证明训练过程保持神经元拓扑不变
- 学习率低于临界值η*时,拓扑结构完全保留;高于时则逐步简化
- 适用于各类架构与损失函数,为深度学习提供通用拓扑分析框架
我们证明,在一大类排列等变学习规则(包括SGD、Adam等)下,训练过程在神经元间诱导出双李普希茨映射,并强烈约束神经元分布的拓扑结构。该结果揭示了小学习率与大学习率 $η$ 的定性差异:当 $η < η^*$ 时,训练严格保持神经元的所有拓扑结构;而当 $η > η^*$ 时,学习过程允许拓扑简化,使神经元流形逐渐粗糙化,从而降低模型表达能力。结合近期发现的‘稳定性边缘’现象,梯度下降下的神经网络学习动态可分为两个阶段:先在拓扑约束下进行平滑优化,随后进入通过剧烈拓扑简化进行学习的第二阶段。本理论不依赖特定架构或损失函数,可普遍应用于深度学习的拓扑分析。
原文摘要 · Abstract (English)
We prove that for a broad class of permutation-equivariant learning rules (including SGD, Adam, and others), the training process induces a bi-Lipschitz mapping between neurons and strongly constrains the topology of the neuron distribution during training. This result reveals a qualitative difference between small and large learning rates $η$. With a learning rate below a topological critical point $η^*$, the training is constrained to preserve all topological structure of the neurons. In contrast, above $η^*$, the learning process allows for topological simplification, making the neuron manifold progressively coarser and thereby reducing the model's expressivity. Viewed in combination with the recent discovery of the edge of stability phenomenon, the learning dynamics of neuron networks under gradient descent can be divided into two phases: first they undergo smooth optimization under topological constraints, and then enter a second phase where they learn through drastic topological simplifications. A key feature of our theory is that it is independent of specific architectures or loss functions, enabling the universal application of topological methods to the study of deep learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。