用分层线性模型理解深度神经网络的动态现象
Position: Solve Layerwise Linear Models First to Understand Neural Dynamical Phenomena (Neural Collapse, Emergence, Lazy/Rich Regime, and Grokking)
- 以分层线性模型为简化框架,揭示神经网络演化核心机制
- 成功解释神经坍缩、涌现、懒惰/丰富状态及磨合等现象
- 适合想快速把握深度学习动态原理的研究者
在物理学中,复杂系统常被简化为保留核心原理的可解模型。在机器学习中,分层线性模型(如线性神经网络)作为神经网络动态的简化表示,遵循动态反馈原则——各层相互调控并放大彼此演化。该原则不仅适用于简化模型,还能有效解释深度神经网络中的多种动态现象,包括神经坍缩、涌现、懒惰与丰富态,以及磨合现象。本文主张采用保留神经动态核心原理的分层线性模型,以加速深度学习科学的发展。
原文摘要 · Abstract (English)
In physics, complex systems are often simplified into minimal, solvable models that retain only the core principles. In machine learning, layerwise linear models (e.g., linear neural networks) act as simplified representations of neural network dynamics. These models follow the dynamical feedback principle, which describes how layers mutually govern and amplify each other's evolution. This principle extends beyond the simplified models, successfully explaining a wide range of dynamical phenomena in deep neural networks, including neural collapse, emergence, lazy and rich regimes, and grokking. In this position paper, we call for the use of layerwise linear models retaining the core principles of neural dynamical phenomena to accelerate the science of deep learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。