揭示了刚性神经微分方程中梯度消失的根源及其不可回避性。
The Vanishing Gradient Problem for Stiff Neural Differential Equations
- 通过分析稳定性函数导数,揭示梯度消失是刚性求解器的普遍特性。
- 证明所有A-稳定方法在高刚性下参数梯度衰减率至少为O(|z|⁻¹)。
- 为训练刚性神经微分方程提供理论警示,适合动态系统建模研究者。
神经微分方程等参数化动力系统依赖于对数值解的参数求导进行梯度优化。在刚性系统中,控制快速衰减模式的参数敏感性在训练过程中趋于消失,导致优化困难。本文表明,这种梯度消失现象并非特定方法的产物,而是所有A-稳定和L-稳定刚性数值积分方案的普遍特征。我们分析了通用刚性积分方案的有理稳定性函数,证明相关参数敏感性由稳定性函数的导数决定,并在大刚性条件下趋近于零。文中给出了常见刚性积分方案的显式公式,详细展示了其机制。最后严格证明:稳定性函数导数的最慢衰减速率为$O(|z|^{-1})$,揭示了根本限制——所有A-稳定时间步进方法不可避免地抑制刚性条件下的参数梯度,成为训练与参数识别的重大障碍。
原文摘要 · Abstract (English)
Gradient-based optimization of neural differential equations and other parameterized dynamical systems fundamentally relies on the ability to differentiate numerical solutions with respect to model parameters. In stiff systems, it has been observed that sensitivities to parameters controlling fast-decaying modes become vanishingly small during training, leading to optimization difficulties. In this paper, we show that this vanishing gradient phenomenon is not an artifact of any particular method, but a universal feature of all A-stable and L-stable stiff numerical integration schemes. We analyze the rational stability function for general stiff integration schemes and demonstrate that the relevant parameter sensitivities, governed by the derivative of the stability function, decay to zero for large stiffness. Explicit formulas for common stiff integration schemes are provided, which illustrate the mechanism in detail. Finally, we rigorously prove that the slowest possible rate of decay for the derivative of the stability function is $O(|z|^{-1})$, revealing a fundamental limitation: all A-stable time-stepping methods inevitably suppress parameter gradients in stiff regimes, posing a significant barrier for training and parameter identification in stiff neural ODEs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。