揭示门控网络学习时间范围的理论极限,解释为何某些模型更省数据。
Learnability Window in Gated Recurrent Neural Networks
- 提出有效学习率包络函数,量化门控与优化器对时序学习的影响
- 发现学习范围随包络衰减速率呈对数、多项式或指数增长
- 慢衰减包络比增加数据更能提升长时序学习能力,适合高效模型设计
本文建立了一种递归神经网络时序可学习性的统计理论,量化了在有限样本量 $N$ 下,基于梯度的学习方法能够恢复滞后依赖结构的最大时间跨度 $/mathcal{H}_N$。理论核心是有效学习率包络 $f(oldsymbol{ heta})$,该函数刻画了门控机制与自适应优化器共同作用下,状态空间动态与参数更新之间的耦合关系。在重尾(α-稳定)波动下,经验平均以 $N^{-1/κ_α}$ 速率集中,其中 $κ_α = α/(α-1)$。包络衰减与统计集中性的相互作用,导出了 $/mathcal{H}_N$ 的显式标度律:根据 $f(oldsymbol{ heta})$ 的衰减规律,出现对数、多项式和指数三类时序学习模式。结果表明,包络衰减速率是决定时序可学习性的关键因素。较慢的包络衰减能扩大 $/mathcal{H}_N$,而重尾波动通过削弱统计集中性压缩其范围。此外,包络几何结构的影响超过数据规模:减缓包络衰减速率带来的 $/mathcal{H}_N$ 提升,优于单纯增加数据量,因此能实现更慢衰减包络的复杂架构,比简单架构更具数据效率。多个门控架构与优化器的实验验证了这些结构性预测。
原文摘要 · Abstract (English)
We develop a statistical theory of temporal learnability in recurrent neural networks, quantifying the maximal temporal horizon $\mathcal{H}_N$ over which gradient-based learning can recover lag-dependent structure at finite sample size $N$. The theory is built on the effective learning rate envelope $f(\ell)$, a function that captures how gating mechanisms and adaptive optimizers jointly shape the coupling between state-space dynamics and parameter updates during Backpropagation Through Time. Under heavy-tailed ($α$-stable) fluctuations, where empirical averages concentrate at rate $N^{-1/κ_α}$ with $κ_α= α/(α-1)$, the interplay between envelope decay and statistical concentration yields explicit scaling laws for the growth of $\mathcal{H}_N$: logarithmic, polynomial, and exponential temporal learning regimes emerge according to the decay law of $f(\ell)$. These results identify envelope decay as the key determinant of temporal learnability. Slower attenuation of $f(\ell)$ enlarges $\mathcal{H}_N$, while heavy-tailed fluctuations compress it by weakening statistical concentration. Moreover, envelope geometry outweighs dataset size: slowing the envelope's decay enlarges $\mathcal{H}_N$ more than adding data, so more complex architectures that realize slower-decaying envelopes can be more data-efficient than simpler ones. Experiments across multiple gated architectures and optimizers corroborate these structural predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。