用物理相变类比深度学习,揭示正则化如何触发特征学习的级联过程。
Cascading Through the Hierarchy: Regularizer-Induced Feature Detection as Phase Transitions in Deep Linear Neural Networks

- 以正则化强度为参数,构建可解析的线性神经网络模型。
- 发现相变级联与损失曲面几何直接相关,可预测其临界点。
- 连接宏观特征学习与微观损失曲面的海森谱,适合理论研究者。
深度学习的科学理论日益受到关注,其核心是可解析的简化模型,用于精确分析学习动态。本文将正则化强度作为可调外部参数,类比统计物理中的外场,对一个可解析的线性神经网络模型进行严格分析。此前研究中,(i) 已解析预测学习过渡的出现,(ii) 数值研究表明调节正则化强度可引发一系列相变级联,其数量与模型复杂度决定的损失曲面几何有关。本文建立严谨框架,揭示这些相变级联、可学习特征与底层几何之间的精确联系。我们提供相变的解析预测及与所学特征相关的可观测序参量。在最简模型层面,将宏观有效描述与微观损失曲面几何(通过海森谱刻画)联系起来。该模型为基于统计物理概念推进深度学习科学理论提供了坚实平台。
原文摘要 · Abstract (English)
A scientific theory of deep learning, comprising learning dynamics and statistical properties of learned models, is rapidly gaining attention. One of the corner stones of this development are analytically solvable toy models, allowing for the fully tractable analysis of the learning dynamics. Here we analytically investigate such a toy model using the regularization strength as a tunable external parameter - akin to external fields in statistical physics. In previous studies, (i) an onset of learning transition was predicted analytically and (ii) it was phenomenologically/numerically established that tuning the regularization strength can result in a cascade of phase transitions. The number of those transitions was linked to the geometry of the loss landscape determined by the model complexity. Setting up a rigorous framework underpinning the previous numerical observations, our investigation reveals a precise connection between those cascades of phase transitions, learnable features and the underlying geometry. We provide analytic predictions of these phase transitions as well as tractable order parameters related to learned features. At the level of the minimal model, we connect this macroscopic perspective (that can be condensed into an effective description) to the microscopic perspective in terms of the geometry of the loss landscape characterized by the Hessian spectrum. Thus, the presented model provides a platform to explore and sharpen advances made in the scientific theory of deep learning rooted in statistical physics concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。