解析了状态空间模型中奇异微分方程的数学基础与数值解法。
Numerical Analysis of HiPPO-LegS ODE for Deep State Space Models
- 证明奇异的HiPPO-LegS微分方程在数学上适定,虽无任意初值自由度。
- 建立针对黎曼可积输入函数的数值离散化方案收敛性理论。
- 为长序列建模的连续时间状态空间模型提供严谨数学支撑,适合研究者参考。
在深度学习中,近期引入的状态空间模型利用HiPPO(高阶多项式投影算子)记忆单元,通过常微分方程(ODE)近似输入函数的连续轨迹,已在捕捉长序列中的长距离依赖方面展现出实证成功。然而,这些ODE的数学基础,尤其是奇异的HiPPO-LegS(勒让德缩放)微分方程及其对应的数值离散化方法,仍不明确。本文填补这一空白:证明尽管存在奇异性,HiPPO-LegS ODE仍是适定的,但不具备任意初始条件的自由度;进一步建立了针对黎曼可积输入函数的数值离散化方案的收敛性。
原文摘要 · Abstract (English)
In deep learning, the recently introduced state space models utilize HiPPO (High-order Polynomial Projection Operators) memory units to approximate continuous-time trajectories of input functions using ordinary differential equations (ODEs), and these techniques have shown empirical success in capturing long-range dependencies in long input sequences. However, the mathematical foundations of these ODEs, particularly the singular HiPPO-LegS (Legendre Scaled) ODE, and their corresponding numerical discretizations remain unsettled. In this work, we fill this gap by establishing that HiPPO-LegS ODE is well-posed despite its singularity, albeit without the freedom of arbitrary initial conditions. Further, we establish convergence of the associated numerical discretization schemes for Riemann integrable input functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。