arXiv:2508.07370cs.LG2025-08中稿 · ICLR被引 3

揭示深度神经网络训练中的内在动态机制,解释参数为何趋向低维结构。

Intrinsic training dynamics of deep neural networks

  • 提出参数映射下的内在梯度流理论,建立训练动态与低维结构的联系
  • 证明在特定初始化下,深层ReLU网络可转化为仅依赖初始状态的低维动态
  • 推广平衡初始化至更广的松弛平衡情形,适用于线性网络极限情况

深度学习理论中的一个根本挑战是理解基于梯度的训练是否能引导参数进入某些低维结构(如稀疏或低秩集合),从而产生隐式偏差。作为基础步骤,受现有隐式偏差分析证明结构的启发,本文研究在参数 $θ$ 上的梯度流何时可对应于一个“提升”变量 $z = ϕ(θ)$ 的内在梯度流,其中 $ϕ$ 为与架构相关的函数。本文定义了一种内在动态性质,并揭示其与因子分解 $ϕ$ 相关守恒律的关系。由此导出一个基于线性映射核包含关系的简单判据,给出该性质成立的必要条件。随后将理论应用于任意深度的ReLU网络,证明在稠密初始值集合上,当 $ϕ$ 为路径提升(path-lifting)时,可将梯度流重写为仅依赖 $z$ 和初始化的低维内在动态。对于线性网络中 $ϕ$ 为权矩阵乘积的情形,已知在平衡初始化下存在内在动态;本文将其推广至更广泛的“松弛平衡”初始化,并在特定配置下证明这些初始化是确保内在度量性质的唯一选择。最后,对无穷深线性网络对应的线性神经微分方程,在松弛平衡初始化下显式给出了相应的内在动态。

原文摘要 · Abstract (English)

A fundamental challenge in the theory of deep learning is to understand whether gradient-based training can promote parameters belonging to certain lower-dimensional structures (e.g., sparse or low-rank sets), leading to so-called implicit bias. As a stepping stone, motivated by the proof structure of existing implicit bias analyses, we study when a gradient flow on a parameter $θ$ implies an intrinsic gradient flow on a ``lifted'' variable $z = ϕ(θ)$, for an architecture-related function $ϕ$. We express a so-called intrinsic dynamic property and show how it is related to the study of conservation laws associated with the factorization $ϕ$. This leads to a simple criterion based on the inclusion of kernels of linear maps, which yields a necessary condition for this property to hold. We then apply our theory to general ReLU networks of arbitrary depth and show that, for a dense set of initializations, it is possible to rewrite the flow as an intrinsic dynamic in a lower dimension that depends only on $z$ and the initialization, when $ϕ$ is the so-called path-lifting. In the case of linear networks with $ϕ$, the product of weight matrices, the intrinsic dynamic is known to hold under so-called balanced initializations; we generalize this to a broader class of {\em relaxed balanced} initializations, showing that, in certain configurations, these are the \emph{only} initializations that ensure the intrinsic metric property. Finally, for the linear neural ODE associated with the limit of infinitely deep linear networks, with relaxed balanced initialization, we make explicit the corresponding intrinsic dynamics.

深度学习理论隐式偏差神经网络动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。