arXiv:2606.29519cs.LGphysics.data-an2026-06

揭示了循环网络中长时依赖学习的动态机制,关键在于时间尺度的自发扩展。

Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks

论文配图:Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks
图 1 · 摘自论文原文
  • 通过耦合状态与参数动力学,发现时间尺度可自发扩展而非由架构固定
  • 当重尾波动足够强时,网络进入慢遗忘的幂律衰减态,学习成本保持多项式增长
  • 适合研究长期记忆机制或设计具备长程学习能力的神经网络的研究者

用随机梯度下降训练的循环网络难以实现长程学习,因为过去输入的影响随滞后量ℓ衰减,若衰减过快则无法从有限数据中学习。这种衰减由包络函数f(ℓ)描述:指数衰减使学习ℓ阶依赖所需数据呈指数增长,而幂律衰减则保持多项式增长。我们证明,f(ℓ)的渐近行为并非由架构决定,而是由状态动力学与参数动力学耦合后涌现的结果——系统会落入坍塌态(快速指数遗忘)或非坍塌态(缓慢幂律遗忘)。其核心是两者间的竞争:训练倾向于缩短有效时间尺度,而学习动力学中的罕见重尾波动则将部分尺度推向极长值。只有当重尾扰动足够强以抵消训练拉力时,非坍塌态才能维持。我们构建粗粒度随机过程,推导出该路径的明确阈值。一个谱指数β同时控制时间尺度分布范围和遗忘速度。实际实现该态还需架构与优化器共同支持宽谱时间尺度生成;若此能力受限,即使施加强重尾驱动,网络仍会坍塌。因此,重尾波动不是需抑制的噪声,而是维系长程学习的关键机制。

原文摘要 · Abstract (English)

Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data. This fade is captured by an envelope $f(\ell)$. An exponential fade makes the data needed to learn a lag-$\ell$ dependence grow exponentially, putting long horizons out of reach; a power-law fade keeps the cost polynomial. We show that the asymptotic decay behavior of $f(\ell)$ is not fixed by the architecture. Instead, it emerges from the coupling between the state dynamics and parameter dynamics, settling into either a collapsed regime (fast, exponential forgetting) or an extended, anti-collapsed regime (slow, power-law forgetting). The intuition is a competition within these coupled dynamics. Training drives the network's effective time scales toward short ones, while rare, heavy-tailed fluctuations of the learning dynamics push a few of them to very long values. Along the route studied here, the extended regime survives only when these heavy-tailed pushes are strong enough to balance the pull. We make this mathematically precise with a coarse-grained stochastic process and derive an explicit threshold at which this route to the extended regime becomes available. A single exponent, the spectral exponent~$β$, then governs both the spread of time scales and how slowly the network forgets. Realizing the regime in practice needs one more ingredient: the joint action of the architecture and the optimizer must be able to hold such a broad spread. A network whose capacity to generate broad time-scale spectra is severely constrained still collapses, even when supplied with strong heavy-tailed forcing. Heavy-tailed fluctuations thus act not as noise to be suppressed, but as the mechanism that sustains long-range learning.

循环网络长时依赖动力系统时间尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。