发现神经网络训练存在敏感探索与稳定优化的两阶段动态。
New Evidence of the Two-Phase Learning Dynamics of Neural Networks
- 通过时间窗分析,揭示训练早期对初始条件敏感的混沌现象。
- 训练中期后,模型函数轨迹被限制在狭窄的锥形区域中。
- 为理解深度学习的渐进式学习机制提供新视角,适合研究训练动力学者。
理解深度神经网络的学习机制仍是现代机器学习中的基础挑战。越来越多的证据表明,训练过程存在显著的相变现象,但对其本质的理解仍不充分。本文提出一种区间化视角,通过比较时间窗口内网络状态的变化,揭示了两个新现象,阐明了深度学习的两阶段特性。第一, extbf{混沌效应}:在不同训练阶段引入极微小的参数扰动,发现网络响应从混沌转为稳定,表明存在一个早期临界期,此时网络对初始条件高度敏感;第二, extbf{锥形效应}:追踪经验神经正切核(eNTK)的演化,发现在该过渡点之后,模型的功能轨迹被限制在一个狭长的锥形子空间内——尽管核仍在变化,却落入一个紧密的角度区域内。这两类效应共同提供了深度网络从敏感探索到稳定精炼的结构性、动态性视图。
原文摘要 · Abstract (English)
Understanding how deep neural networks learn remains a fundamental challenge in modern machine learning. A growing body of evidence suggests that training dynamics undergo a distinct phase transition, yet our understanding of this transition is still incomplete. In this paper, we introduce an interval-wise perspective that compares network states across a time window, revealing two new phenomena that illuminate the two-phase nature of deep learning. i) \textbf{The Chaos Effect.} By injecting an imperceptibly small parameter perturbation at various stages, we show that the response of the network to the perturbation exhibits a transition from chaotic to stable, suggesting there is an early critical period where the network is highly sensitive to initial conditions; ii) \textbf{The Cone Effect.} Tracking the evolution of the empirical Neural Tangent Kernel (eNTK), we find that after this transition point the model's functional trajectory is confined to a narrow cone-shaped subset: while the kernel continues to change, it gets trapped into a tight angular region. Together, these effects provide a structural, dynamical view of how deep networks transition from sensitive exploration to stable refinement during training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。