arXiv:2506.08764cs.LG2025-06被引 4

揭示深度网络雅可比矩阵稳定性新理论,支持稀疏与相关权重。

On the Stability of the Jacobian Matrix in Deep Neural Networks

  • 基于随机矩阵论,推导出适用于稀疏和相关权重的通用稳定性定理。
  • 证明在结构化权重下,网络梯度仍可避免爆炸或消失。
  • 为现代神经网络初始化提供更广泛理论支撑,适合研究者参考。

深度神经网络随深度增加常出现梯度爆炸或消失问题,这与输入输出雅可比矩阵的谱特性密切相关。以往研究仅限于权重独立同分布的全连接网络,而本文突破这一限制:基于随机矩阵论最新进展,建立适用于稀疏(如剪枝引入)及非独立同分布、弱相关权重(如训练导致)的深层网络通用稳定性定理。该结果为具有结构化和依赖随机性的现代神经网络提供了严格的谱稳定性保证,拓展了初始化方案的理论基础。

原文摘要 · Abstract (English)

Deep neural networks are known to suffer from exploding or vanishing gradients as depth increases, a phenomenon closely tied to the spectral behavior of the input-output Jacobian. Prior work has identified critical initialization schemes that ensure Jacobian stability, but these analyses are typically restricted to fully connected networks with i.i.d. weights. In this work, we go significantly beyond these limitations: we establish a general stability theorem for deep neural networks that accommodates sparsity (such as that introduced by pruning) and non-i.i.d., weakly correlated weights (e.g. induced by training). Our results rely on recent advances in random matrix theory, and provide rigorous guarantees for spectral stability in a much broader class of network models. This extends the theoretical foundation for initialization schemes in modern neural networks with structured and dependent randomness.

深度学习雅可比矩阵稳定性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。