arXiv:2512.18209cs.LGcs.AI2025-12被引 1

揭示深度学习中幂律谱动态的理论条件,明确何时训练会呈现幂律行为。

When Does Learning Renormalize? Sufficient Conditions for Power Law Spectral Dynamics

  • 基于广义分辨率壳动力学框架,构建多层约束条件
  • 发现幂律行为源于对称性与时间缩放的耦合效应
  • 为理解深层网络训练动力学提供可验证的理论依据

现代深度学习系统中广泛观测到的幂律尺度现象,其理论起源和适用范围仍不清晰。广义分辨率壳动力学(GRSD)框架将学习过程建模为对数分辨率壳间的谱能量传输,提供了训练的动力学粗粒化描述。在该框架中,幂律尺度对应于一种简化的可重整化壳动力学,但这种行为并非必然出现,需学习过程具备额外结构特性。本文识别出一组充分条件,使GRSD壳动力学具有可重整化的粗粒化描述:包括计算图中梯度传播的有界性、初始化时弱函数非相干性、训练过程中雅可比矩阵的受控演化,以及重整化壳耦合的对数平移不变性。进一步证明,幂律尺度并非仅由可重整性决定,而是刚性结果:一旦对数平移不变性与梯度流的内在时间缩放协变性结合,重整化后的GRSD速度场即被强制为幂律形式。

原文摘要 · Abstract (English)

Empirical power--law scaling has been widely observed across modern deep learning systems, yet its theoretical origins and scope of validity remain incompletely understood. The Generalized Resolution--Shell Dynamics (GRSD) framework models learning as spectral energy transport across logarithmic resolution shells, providing a coarse--grained dynamical description of training. Within GRSD, power--law scaling corresponds to a particularly simple renormalized shell dynamics; however, such behavior is not automatic and requires additional structural properties of the learning process. In this work, we identify a set of sufficient conditions under which the GRSD shell dynamics admits a renormalizable coarse--grained description. These conditions constrain the learning configuration at multiple levels, including boundedness of gradient propagation in the computation graph, weak functional incoherence at initialization, controlled Jacobian evolution along training, and log--shift invariance of renormalized shell couplings. We further show that power--law scaling does not follow from renormalizability alone, but instead arises as a rigidity consequence: once log--shift invariance is combined with the intrinsic time--rescaling covariance of gradient flow, the renormalized GRSD velocity field is forced into a power--law form.

深度学习动力学幂律可重整化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。