arXiv:2605.12763cs.LGmath.DS2026-05

揭示了循环网络在动态突变时的学习机制,发现学习集中在低秩通道。

Center-Manifold Reduction of Learning at Bifurcations: Interference and Rich Learning in Recurrent Neural Networks

论文配图:Center-Manifold Reduction of Learning at Bifurcations: Interference and Rich Learning in Recurrent Neural Networks
图 1 · 摘自论文原文
  • 基于中心流形理论,将参数到状态的雅可比矩阵简化为低秩算子。
  • 全局神经正切核与费舍尔信息矩阵在突变点被强烈放大并呈各向异性。
  • 适用于研究复杂任务干扰与突变学习现象,尤其适合深度循环网络研究者。

循环神经网络中的丰富学习常通过潜在动力学的突然转变实现,但目前缺乏理论预测梯度下降在此类事件中的行为。本文通过全局经验神经正切核(GeNTK)研究了一阶分歧点附近的局部学习几何。在中心流形条件下,当分歧相关敏感性主导有界残差项时,证明参数到状态的雅可比矩阵 $D_θh$ 可近似为一个低秩正规型算子。由此产生的 GeNTK 和费舍尔信息矩阵因此显著放大且高度各向异性,分别集中于四类标量一阶分歧的秩一通道和尼马克-萨克尔分歧的秩二实通道。高维 RNN 的受控实验验证了该算子降维。在实际学习的 RNN 中,相同的低秩集中对应于损失的突变和子任务干扰;局部投影能预测这些效应的符号。此外,在输入驱动的 15 任务漏失率 RNN(LeakyRNN)中,GeNTK 放大与记忆动力学(MemoryPro)的延续检测变化一致。结果表明,可在算子层面描述动态突变附近的学习,同时提供可扩展的低维学习几何诊断方法。

原文摘要 · Abstract (English)

Rich learning in recurrent neural networks often proceeds through sudden transitions in latent dynamics, but there is little theory predicting how gradient descent behaves during these events. We study the local learning geometry near codimension-one bifurcations through the global empirical Neural Tangent Kernel (GeNTK). Under local center-manifold conditions, and when bifurcation-related sensitivity dominates bounded residual terms, we show that the global parameter-to-state Jacobian \(D_θh\) is approximated by a low-rank normal-form operator. The induced GeNTK and Fisher information matrix therefore become strongly amplified and anisotropic, concentrating toward a rank-one channel for the four scalar codimension-one bifurcations and a rank-two real channel for a Neimark--Sacker bifurcation. Controlled high-dimensional RNN experiments validate this operator reduction. In learned RNNs, the same low-rank concentration coincides with abrupt loss changes and subtask interference, while a local projection predicts the sign of these effects near isolated events. Finally, in an input-driven 15-task LeakyRNN, GeNTK amplification aligns with continuation-detected changes in the MemoryPro dynamics. These results suggest a tractable operator-level description of learning near dynamical transitions, together with scalable diagnostics for amplified low-dimensional learning geometry.

循环网络动态突变学习几何低秩结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。