arXiv:2501.02378cs.LGq-bio.NC2025-01被引 1

揭示了循环网络突变学习的内在机制,解释为何训练会突然失败或成功。

A ghost mechanism: An analytical model of abrupt learning in recurrent networks

  • 提出‘幽灵机制’模型,用单参数控制动态系统中的瞬时减速现象。
  • 发现学习率超过临界值时,梯度消失与振荡导致学习崩溃。
  • 适用于低秩和全秩网络,为改善训练提供稳定化方法。

在工作记忆任务中,循环神经网络(RNNs)常出现突变学习现象。本文提出‘幽灵机制’,描述动力系统在鞍结分岔残余点附近的瞬时减速行为。通过降维,建立一维规范形式,以单一尺度参数解析学习过程。研究发现,当学习率超过临界值(反比于计算时长)时,学习因两种模式崩溃:(i) 梯度消失;(ii) 最小值附近梯度振荡。此时系统可能陷入高置信但错误预测的‘无学习区’,即参数空间中梯度消失区域。在低秩RNN中验证了幽灵点出现在突变前,并在全秩RNN的典型工作记忆任务中证明其普遍性。理论建议:提升可训练秩可稳定学习轨迹,降低输出置信度可缓解困于无学习区。整体表明,任务计算需求限制优化景观,导致常见学习困难部分源于需实现的动力学结构。

原文摘要 · Abstract (English)

Abrupt learning is a common phenomenon in recurrent neural networks (RNNs) trained on working memory tasks. In such cases, the networks develop transient slow regions in state space that extend the effective timescales of computation. However, the mechanisms driving sudden performance improvements and their causal role remain unclear. To address this gap, we introduce the ghost mechanism, a process by which dynamical systems exhibit transient slowdown near the remnant of a saddle-node bifurcation. By reducing the high-dimensional dynamics near ghost points, we derive a one-dimensional canonical form that analytically captures learning as a process controlled by a single scale parameter. Using this model, we study a form of abrupt learning emerging from ghost points and identify a critical learning rate that scales as an inverse power law with the timescale of the learned computation. Beyond this rate, learning collapses through two interacting modes: (i) vanishing gradients and (ii) oscillatory gradients near minima. These features can lock the system into high-confidence but incorrect predictions when parameter updates trigger a no-learning zone, a region of parameter space where gradients vanish. We validate these predictions in low-rank RNNs, where ghost points precede abrupt transitions, and further demonstrate their generality in full-rank RNNs trained on canonical working memory tasks. Our theory offers two approaches to address these learning difficulties: increasing trainable ranks stabilizes learning trajectories, while reducing output confidence mitigates entrapment in no-learning zones. Overall, the ghost mechanism reveals how the computational demands of a task constrain the optimization landscape, demonstrating that well-known learning difficulties in RNNs partly arise from the dynamical systems they must learn to implement.

循环网络学习机制动力系统梯度问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。