arXiv:2507.06381cs.LGcs.AI2025-07被引 3

揭示梯度下降中循环网络动态坍缩的数学机制,提出新分析工具。

KPFlow: An Operator Perspective on Dynamic Collapse Under Gradient Descent Training of Recurrent Networks

  • 将梯度流分解为参数算子与线性化传播算子的乘积,揭示动态演化机理。
  • 发现低维潜在动态源于网络结构本身,而非任务特性,解释坍缩本质。
  • 提供开源工具包KPFlow,支持多任务训练中子任务目标对齐分析。

梯度下降及其变体是训练循环动力系统(如循环神经网络、神经微分方程和门控循环单元)的主要方法。这些模型的动态特性表现出神经坍缩和潜在表征涌现等现象,可能支撑其优异的泛化能力。在神经科学中,这些表征特征被用于对比生物与人工系统的学习过程。尽管已有进展,仍缺乏理论工具来严谨理解有限非线性模型中所学表征的形成机制。本文证明,描述模型动态演化的梯度流可分解为两个算子的乘积:参数算子K和线性化流传播算子P。K类似于前馈网络中的神经正切核,而P出现在李雅普诺夫稳定性与最优控制理论中。我们展示了该分解的两个应用:第一,揭示二者相互作用如何导致梯度下降下的低维潜在动态,并说明坍缩由网络结构决定,超出任务本身的性质;第二,在多任务训练中,利用算子衡量各子任务相关目标的对齐程度。通过实验与理论验证,我们推出了开源工具包KPFlow,实现对一般循环架构的稳健分析。本工作推动了对非线性循环模型中梯度下降学习机制的深层理解。

原文摘要 · Abstract (English)

Gradient Descent (GD) and its variants are the primary tool for enabling efficient training of recurrent dynamical systems such as Recurrent Neural Networks (RNNs), Neural ODEs and Gated Recurrent units (GRUs). The dynamics that are formed in these models exhibit features such as neural collapse and emergence of latent representations that may support the remarkable generalization properties of networks. In neuroscience, qualitative features of these representations are used to compare learning in biological and artificial systems. Despite recent progress, there remains a need for theoretical tools to rigorously understand the mechanisms shaping learned representations, especially in finite, non-linear models. Here, we show that the gradient flow, which describes how the model's dynamics evolve over GD, can be decomposed into a product that involves two operators: a Parameter Operator, K, and a Linearized Flow Propagator, P. K mirrors the Neural Tangent Kernel in feed-forward neural networks, while P appears in Lyapunov stability and optimal control theory. We demonstrate two applications of our decomposition. First, we show how their interplay gives rise to low-dimensional latent dynamics under GD, and, specifically, how the collapse is a result of the network structure, over and above the nature of the underlying task. Second, for multi-task training, we show that the operators can be used to measure how objectives relevant to individual sub-tasks align. We experimentally and theoretically validate these findings, providing an efficient Pytorch package, \emph{KPFlow}, implementing robust analysis tools for general recurrent architectures. Taken together, our work moves towards building a next stage of understanding of GD learning in non-linear recurrent models.

循环网络梯度下降动态坍缩算子分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。