解析深度线性网络中初始化对学习动态的影响,揭示从丰富到懒惰的完整演化机制。
From Lazy to Rich: Exact Learning Dynamics in Deep Linear Networks
- 通过λ平衡初始化推导出精确解,刻画不同初始化下的学习过程
- 发现权重尺度分布决定表示是否动态演化,影响模型性能上限
- 为持续学习、迁移学习等场景提供理论依据,适合研究模型训练机制者
生物和人工神经网络会发展出内部表征以完成复杂任务。在人工网络中,这些模型的有效性依赖于构建任务相关表征的能力,这一过程受数据集、架构、初始化策略与优化算法之间相互作用的影响。已有研究指出,不同的初始化会使网络处于‘懒惰’(表示保持静态)或‘丰富’(表示动态演化)的学习模式。本文研究了深度线性网络中初始化如何影响学习动态,推导出由层间权重相对尺度定义的λ平衡初始化下的精确解。这些解捕捉了从丰富到懒惰模式全谱系中表示的演化过程以及神经正切核(NTK)的变化。研究深化了对权重初始化如何影响学习模式的理论理解,对持续学习、逆向学习及迁移学习具有启示意义,适用于神经科学与实际应用。
原文摘要 · Abstract (English)
Biological and artificial neural networks develop internal representations that enable them to perform complex tasks. In artificial networks, the effectiveness of these models relies on their ability to build task specific representation, a process influenced by interactions among datasets, architectures, initialization strategies, and optimization algorithms. Prior studies highlight that different initializations can place networks in either a lazy regime, where representations remain static, or a rich/feature learning regime, where representations evolve dynamically. Here, we examine how initialization influences learning dynamics in deep linear neural networks, deriving exact solutions for lambda-balanced initializations-defined by the relative scale of weights across layers. These solutions capture the evolution of representations and the Neural Tangent Kernel across the spectrum from the rich to the lazy regimes. Our findings deepen the theoretical understanding of the impact of weight initialization on learning regimes, with implications for continual learning, reversal learning, and transfer learning, relevant to both neuroscience and practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。