揭示深度网络中权重诱导的层间度量层次结构
Learning Dynamics Reveal a Hierarchy of Weight-Induced Layerwise Gram Metrics
- 用激活场与共轭场描述梯度下降的集体动力学
- 三层及以上网络首次出现权重诱导的拉回度量
- 适合研究神经网络动态与几何结构的学者
我们研究具有固定读出和二次损失的前馈ReLU网络,将梯度下降重写为训练集上激活场与共轭场的集体动力学。在学习率的一阶近似下,推导出单、双、三层的情况,并给出任意深度的递推关系。单层时激活动力学直接闭合,残差更新由输入格拉姆矩阵与共激活/反向传播格拉姆矩阵的乘积决定;两层需引入共轭场,但尚未出现非平凡的拉回格拉姆度量;三层时首个权重诱导的拉回格拉姆度量进入共轭场动力学。任意深度下,激活变化通过递归响应算子 \\cU_\ell^{αβ} 前向传播,共轭场变化通过有效传输算子 \\cM_\ell^{αβ} 反向传播。它们的收缩重构出层间残差核 \\ K_{αβ}^{(L)}=\sum_{\ell=1}^{L}Q_{αβ}^{(\ell-1)}S_{αβ}^{(\ell)}。该描述揭示每层截面处前推与拉回传输的对偶性,并将格拉姆度量识别为更广泛激活条件传输算子中的最低阶非平凡项。
原文摘要 · Abstract (English)
We study feed-forward ReLU networks with fixed readout and quadratic loss, and rewrite gradient descent as a collective dynamics of activation fields and conjugate fields on the training set. Working to first order in the learning rate inside a fixed activation chamber, we derive explicitly the one-, two- and three-hidden-layer cases, and then give the arbitrary-depth recursion. For one hidden layer the activation dynamics closes directly and the residual update is governed by the product of an input Gram matrix and a co-activation/backpropagation Gram matrix. For two hidden layers a conjugate field is required, but no nontrivial pullback Gram metric has yet appeared. For three hidden layers the first weight-induced pullback Gram metric enters the conjugate-field dynamics. At arbitrary depth, activation variations propagate forward through a recursive response operator \(\cU_\ell^{αβ}\), while conjugate-field variations propagate backward through an effective transport operator \(\cM_\ell^{αβ}\). Their contractions reconstruct a layerwise residual kernel \[ K_{αβ}^{(L)}=\sum_{\ell=1}^{L}Q_{αβ}^{(\ell-1)}S_{αβ}^{(\ell)}. \] The resulting description exposes a duality between push-forward and pullback transport across every layer cut, and identifies the first Gram metrics as the lowest nontrivial terms in a broader hierarchy of activation-conditioned transport operators. We deliberately stop at the level of collective fields, conjugate fields, residual kernels and cut-wise transport metrics, leaving the later tensorial geometric formulation outside the scope of this paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。