arXiv:2602.05779cs.LGcs.IT2026-02被引 1

通过调控方差提升稀疏激活网络训练稳定性,实现高达90%的稀疏度。

How Controlling the Variance can Improve Training Stability of Sparsely Activated DNNs and CNNs

  • 提出新初始化策略,增大稀疏激活下的高斯过程方差以增强稳定性。
  • 实验证明可支持深度网络隐藏层达90%稀疏度,且训练更稳定。
  • 适用于追求模型轻量化与高效推理的深度神经网络研究者。

针对深度网络随机初始化的边缘混沌(EoC)理论,通过将中间层表征为高斯过程,既能保持初始输出信息,又能最小化梯度爆炸或消失问题,从而提供权重和偏置初始化方差的计算公式。对于近似线性激活函数,该理论通常建议随着网络深度增加,高斯过程方差趋于零。本文研究了较少被关注的高稀疏性激活情形——在原点附近大量值被置零。在此设定下,我们证明了一种新现象:更大的固定高斯过程方差反而有助于训练稳定性。这一发现指导我们提出一种简单但有效的初始化方法,使深度神经网络(DNNs)和卷积神经网络(CNNs)在隐藏层实现高达90%的稀疏度时仍能稳定训练。

原文摘要 · Abstract (English)

The Edge-of-Chaos (EoC) theory developed for the random initialization of deep networks allows more efficient training by both preserving information in the initial outputs of the network and minimising exploding or vanishing gradients through characterisation of the intermediate layers as Gaussian processes. This EoC theory provides formulae for the choice of the initialisation distribution variances of the weights and biases. For activations which are approximately linear around the origin, the EoC theory typically encourages the Gaussian process variance to converge towards zero with increasing depth. Here we consider the less studied setting of highly sparsity inducing activations where a large region of values near the origin are set to zero. In this setting we prove a new phenomenon whereby initialisations leading to larger fixed Gaussian processes are beneficial to training stability. This theory informs a new, yet simple, initialisation strategy that allows training DNNs and CNNs with as large as 90\% sparsity in the hidden layers.

稀疏网络初始化策略训练稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。