揭示深度非线性网络中梯度逃逸的临界规律,解释训练为何分阶段跃迁。
A Theory of Saddle Escape in Deep Nonlinear Networks
- 基于弗罗贝尼乌斯范数不平衡恒等式,将复杂网络降维为标量微分方程。
- 逃逸时间遵循 τ⋆ = Θ(ε^-(r-2)),由瓶颈层数量 r 决定而非总层数。
- 理论与数值模拟高度吻合,适用于小初始化的深层网络研究者。
在小初始化的深度网络中,训练过程呈现长平台期,其间穿插急剧的特征获取跃迁。尽管浅层非线性网络和深层线性网络已有深入研究,但将其分析拓展至深层非线性网络仍具挑战。本文推导出任意光滑激活函数与可微损失下,各层权重矩阵弗罗贝尼乌斯范数不平衡的精确恒等式,并据此将激活函数划分为四类普适类。在置换对称子流形上,该恒等式与近似平衡律结合,将完整矩阵流简化为标量常微分方程,得出临界逃逸时间律 τ⋆ = Θ(ε^-(r-2)),其指数由瓶颈尺度的层数量 r 决定,而非总深度 L。在 He 正则初始化下,当瓶颈层数量为 r 且按 ε 缩放时,该相同指数 (r-2) 亦被恢复,此时对称子流形虽被流保持但非吸引。理论预测与数值模拟结果高度一致。
原文摘要 · Abstract (English)
In deep networks with small initialization, training exhibits long plateaus separated by sharp feature-acquisition transitions. Whereas shallow nonlinear networks and deep linear networks are well studied, extending these analyses to deep nonlinear networks remains challenging. We derive an exact identity for the imbalance of Frobenius norms of layer weight matrices that holds for any smooth activation and any differentiable loss and use this to classify activation functions into four universality classes. On the permutation-symmetric submanifold, the identity combines with an approximate balance law to reduce the full matrix flow to a scalar ODE, giving a critical-depth escape time law $τ_\star = Θ(\varepsilon^{-(r-2)})$ governed by the number $r$ of layers at the bottleneck scale rather than the total depth $L$. We find that this same $r-2$ exponent is recovered under He-normal initialization with $r$ bottleneck layers rescaled by $\varepsilon$, where the symmetry manifold is preserved by the flow but not attracting. We find close agreement between our theory and numerical simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。