arXiv:2602.10949stat.MLcs.LG2026-02

提出稳定深层网络激活的吕普诺夫初始化方法

Optimal Initialization in Depth: Lyapunov Initialization and Limit Theorems for Deep Leaky ReLU Networks

  • 基于李雅普诺夫指数理论分析深层漏失ReLU网络的激活行为
  • 发现标准初始化在浅宽网络中仍会导致激活消失或爆炸
  • 新方法直接设李雅普诺夫指数为零,提升训练稳定性

深度无偏置随机漏失ReLU网络的有效初始化需要对随机神经网络有深刻理解。本文对深层无偏置随机漏失ReLU网络进行了严格的概率分析,证明了网络激活范数对数的大数定律与中心极限定理,揭示当层数增加时,其增长由一个称为李雅普诺夫指数的参数决定。该参数刻画了激活消失与爆炸之间的尖锐相变,我们显式计算了高斯或正交权重矩阵下的李雅普诺夫指数。结果表明,标准初始化(如He初始化或正交初始化)无法保证低宽深层网络的激活稳定性。基于此理论洞察,我们提出一种新型初始化方法——吕普诺夫初始化,将李雅普诺夫指数设为零,从而实现网络最大可能的稳定性,并在实验中展现出更优的学习性能。

原文摘要 · Abstract (English)

Effective initialization in deep networks requires an understanding of random neural networks. In this work, a rigorous probabilistic analysis of deep bias-free random Leaky ReLU networks is provided. We prove a Law of Large Numbers and a Central Limit Theorem for the logarithm of the norm of network activations, establishing that, as the number of layers increases, their growth is governed by a parameter called the Lyapunov exponent. This parameter characterizes a sharp phase transition between vanishing and exploding activations, and we calculate the Lyapunov exponent explicitly for Gaussian or orthogonal weight matrices. Our results reveal that standard methods, such as He initialization or orthogonal initialization, do not guarantee activation stability for deep networks of low width. Based on these theoretical insights, we propose a novel initialization method, referred to as Lyapunov initialization, which sets the Lyapunov exponent to zero and thereby ensures that the neural network is as stable as possible, leading empirically to improved learning.

深度学习初始化理论分析网络稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。