arXiv:2605.06563cs.LGhep-th2026-05

解释正交初始化如何让有限宽度网络在深层下保持稳定。

Criticality and Saturation in Orthogonal Neural Networks

  • 推导出正交初始化下网络统计量的逐层递推公式。
  • 证明深层网络的有限宽度张量趋于稳定,与激活函数无关。
  • 理论结果经蒙特卡洛实验验证,精度极高。

长期以来已知,将权重矩阵初始化为正交而非独立同分布高斯,可提升训练性能。该现象可通过有限宽度修正来分析,即在无限宽度统计基础上补充关于 $1/ ext{width}$ 的幂级数。近期实验发现,正交初始化网络的张量在深度增大时趋于稳定,而独立同分布初始化则不然。本文推导出正交初始化下网络统计量有限宽度展开中张量的显式逐层递推关系,并扩展了适用于所有 $1/ ext{width}$ 阶次的费曼图方法。进一步证明,这些递推关系能准确再现激活函数具有零不动点时的稳定性现象。本工作首次从理论上解释了正交初始化下有限宽度非线性网络的深层稳定性,填补了文献空白。通过数值求解递推关系及其大深度解析展开,与网络集合的蒙特卡洛估计高度吻合,验证了理论正确性。

原文摘要 · Abstract (English)

It has been known for a long time that initializing weight matrices to be orthogonal instead of having i.i.d. Gaussian components can improve training performance. This phenomenon can be analyzed using finite-width corrections, where the infinite-width statistics are supplemented by a power series in $1/\mathrm{width}$. In particular, recent empirical results by Day et al. show that the tensors appearing in this treatment stabilize for large depth, as opposed to the tensors of i.i.d.-initialized networks. In this article, we derive explicit layer-wise recursion relations for the tensors appearing in the finite-width expansion of the network statistics in the case of orthogonal initializations. We also provide an extension of recently-introduced Feynman diagrams for the corresponding recursions in the i.i.d.-case which are valid to all orders in $1/\mathrm{width}$. Finally, we show explicitly that the recursions we derive reproduce the stability of the finite-width tensors which was observed for activation functions with vanishing fixed point. This work therefore provides a theoretical explanation for the stability of nonlinear networks of finite width initialized with orthogonal weights, closing a long-standing gap in the literature. We validate our theoretical results experimentally by showing that numerical solutions of our recursion relations and their analytical large-depth expansions agree excellently with Monte-Carlo estimates from network ensembles.

神经网络正交初始化深度学习理论递推关系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。