arXiv:2508.02882cs.LGcs.NA2025-08被引 1

通过控制梯度矩阵正交性,实现超深网络的稳定训练

Deep Network Trainability via Persistent Subspace Orthogonality

  • 设计新架构,使网络雅可比矩阵在子空间上保持正交
  • 实验证明该方法可有效抑制梯度消失/爆炸,支持深层网络训练
  • 适合研究深度网络优化与梯度传播机制的研究者

通过反向传播训练神经网络常受梯度消失或爆炸问题困扰。本文通过分析和控制网络雅可比矩阵,设计新型架构以缓解此类问题。我们首次统一刻画了一类具有正交雅可比矩阵的网络结构,涵盖已有架构并导出新的可训练设计。进一步提出‘持续子空间正交性’的宽松概念,适用于雅可比矩阵仅在非平凡子空间上保持等距的更广泛网络。我们提出实用机制以强制满足该条件,并实证表明该性质对充分保持反向传播中的梯度范数至关重要,从而实现超深网络的有效训练。理论得到大量实验支持。

原文摘要 · Abstract (English)

Training neural networks via backpropagation is often hindered by vanishing or exploding gradients. In this work, we design architectures that mitigate these issues by analyzing and controlling the network Jacobian. We first provide a unified characterization for a class of networks with orthogonal Jacobian including known architectures and yielding new trainable designs. We then introduce the relaxed notion of persistent subspace orthogonality. This applies to a broader class of networks whose Jacobians are isometries only on a non-trivial subspace. We propose practical mechanisms to enforce this condition and empirically show that it is necessary to sufficiently preserve the gradient norms during backpropagation, enabling the training of very deep networks. We support our theory with extensive experiments.

深度网络梯度传播正交性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。