用物理模型解释持续学习中遗忘与突现的机制。
Non-Equilibrium Stochastic Dynamics as a Unified Framework for Insight and Repetitive Learning: A Kramers Escape Approach to Continual Learning
- 将神经网络学习建模为非平衡随机系统,通过势能阱和温度变化描述状态演化。
- 证明了弹性权重保持的惩罚项导致过渡速率指数下降,解释了知识遗忘的原理。
- 揭示顿悟与重复练习对应不同温度策略,为智能系统设计噪声调度提供依据。
人工神经网络的持续学习受限于稳定性-可塑性困境:保留旧知识的系统难以学习新知识,反之亦然。现有方法如弹性权重保持(EWC)虽有效,但缺乏对可塑性随任务累积而崩溃的物理解释。此外,顿悟与渐进练习之间的区别也缺乏统一理论描述。本文提出,二者均可在非平衡统计物理框架下统一解释。将学习系统状态视为在双阱势能景观上受朗之万动力学驱动的粒子,其噪声强度由时变有效温度 $T(t)$ 控制。概率密度满足福克-普朗克方程,态间跃迁由克兰默斯逃逸率 $k = (ω_0ω_b/2π)"e^{-ΔE/T}$ 决定。本文贡献有二:第一,识别出 EWC 惩罚项对应能量壁垒,其高度随累积任务数线性增长,导致跃迁速率指数衰减,分析与数值模拟一致;第二,发现顿悟对应 $T(t)$ 的瞬时峰值,驱动快速越障;而重复练习则在适度升高且恒定的温度下,通过持续随机扩散实现转变。该框架为持续学习中的可塑性及其失效提供了物理基础,并建议了人工智能中自适应噪声调度的设计准则。
原文摘要 · Abstract (English)
Continual learning in artificial neural networks is fundamentally limited by the stability--plasticity dilemma: systems that retain prior knowledge tend to resist acquiring new knowledge, and vice versa. Existing approaches, most notably elastic weight consolidation~(EWC), address this empirically without a physical account of why plasticity eventually collapses as tasks accumulate. Separately, the distinction between sudden insight and gradual skill acquisition through repetitive practice has lacked a unified theoretical description. Here, we show that both problems admit a common resolution within non-equilibrium statistical physics. We model the state of a learning system as a particle evolving under Langevin dynamics on a double-well energy landscape, with the noise amplitude governed by a time-dependent effective temperature $T(t)$. The probability density obeys a Fokker--Planck equation, and transitions between metastable states are governed by the Kramers escape rate $k = (ω_0ω_b/2π)\,e^{-ΔE/T}$. We make two contributions. First, we identify the EWC penalty term as an energy barrier whose height grows linearly with the number of accumulated tasks, yielding an exponential collapse of the transition rate predicted analytically and confirmed numerically. Second, we show that insight and repetitive learning correspond to two qualitatively distinct temperature protocols within the same Fokker--Planck equation: insight events produce transient spikes in $T(t)$ that drive rapid barrier crossing, whereas repetitive practice operates at a modestly elevated but fixed temperature, achieving transitions through sustained stochastic diffusion. These results establish a physically grounded framework for understanding plasticity and its failure in continual learning systems, and suggest principled design criteria for adaptive noise schedules in artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。