用谱壳动力学解释神经网络的缩放规律与双下降现象
Renormalizable Spectral-Shell Dynamics as the Origin of Neural Scaling Laws
- 基于函数空间梯度下降,构建谱壳演化方程
- 揭示误差能量在不同频段间流动的自相似规律
- 统一了懒惰学习与特征学习两种训练模式
神经网络的缩放规律和双下降现象表明,尽管深度网络优化过程高度非线性,但仍遵循简单的宏观结构。我们直接从函数空间的梯度下降推导出这一结构。对于均方误差损失,训练误差演化满足 $\\(dot e_t = -M(t)e_t$,其中 $M(t) = J_{θ(t)}J_{θ(t)}^{\\*}$ 是由网络雅可比矩阵诱导的时变自伴算子。利用卡托摄动理论,我们在 $M(t)$ 的瞬时本征基下得到一组精确的耦合模态常微分方程。为提取宏观行为,引入对数谱壳粗粒化,追踪各壳层间的二次误差能量。每个壳层内的微观相互作用在能量层面完全抵消,因此壳层能量仅通过耗散和跨壳层外部相互作用演化。我们提出‘可重整化壳动力学’假设,使得累积微观效应简化为壳边界上的可控通量。在相关分辨率范围内假设有效幂律谱传输,壳动力学具有移动分辨率前沿的自相似解,且显式给出缩放指数。该框架解释了神经网络的缩放规律和双下降,并将懒惰(NTK类)训练与特征学习统一为同一谱壳动力学的两个极限。
原文摘要 · Abstract (English)
Neural scaling laws and double-descent phenomena suggest that deep-network training obeys a simple macroscopic structure despite highly nonlinear optimization dynamics. We derive such structure directly from gradient descent in function space. For mean-squared error loss, the training error evolves as $\dot e_t=-M(t)e_t$ with $M(t)=J_{θ(t)}J_{θ(t)}^{\!*}$, a time-dependent self-adjoint operator induced by the network Jacobian. Using Kato perturbation theory, we obtain an exact system of coupled modewise ODEs in the instantaneous eigenbasis of $M(t)$. To extract macroscopic behavior, we introduce a logarithmic spectral-shell coarse-graining and track quadratic error energy across shells. Microscopic interactions within each shell cancel identically at the energy level, so shell energies evolve only through dissipation and external inter-shell interactions. We formalize this via a \emph{renormalizable shell-dynamics} assumption, under which cumulative microscopic effects reduce to a controlled net flux across shell boundaries. Assuming an effective power-law spectral transport in a relevant resolution range, the shell dynamics admits a self-similar solution with a moving resolution frontier and explicit scaling exponents. This framework explains neural scaling laws and double descent, and unifies lazy (NTK-like) training and feature learning as two limits of the same spectral-shell dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。