arXiv:2505.07244math.DScs.LG2025-05被引 1

研究神经延迟微分方程的记忆容量如何影响其逼近能力。

The Influence of the Memory Capacity of Neural DDEs on the Universal Approximation Property

  • 通过控制李普希茨常数与延迟的乘积来调节记忆容量。
  • 记忆容量不足时无法实现通用逼近,足够大时可实现。
  • 适合研究连续动力系统建模与深度网络理论的学者。

神经常微分方程(Neural ODEs)是残差网络(ResNets)的连续时间版本,而神经延迟微分方程(Neural DDEs)可视为密集连接残差网络(DenseResNets)的无限深度极限。与传统ResNet不同,DenseResNets允许跨层的捷径连接,引入了网络中的记忆特性。本文研究记忆容量对神经DDE通用逼近性质的影响,关键参数为李普希茨常数与延迟的乘积 $Kτ$。在非增广架构下(宽度不超过输入输出维度),若 $Kτ$ 足够小,神经DDE动态可被神经ODE逼近,因此也缺乏通用逼近能力;若 $Kτ$ 足够大,则神经DDE对连续函数具备通用逼近性。若采用增广架构,可扩大实现通用逼近的参数范围。结果表明,仅当记忆容量超过某一阈值时,通用逼近才能成立,而非仅依赖于正延迟的无限维相空间。

原文摘要 · Abstract (English)

Neural Ordinary Differential Equations (Neural ODEs), which are the continuous-time analog of Residual Neural Networks (ResNets), have gained significant attention in recent years. Similarly, Neural Delay Differential Equations (Neural DDEs) can be interpreted as an infinite depth limit of Densely Connected Residual Neural Networks (DenseResNets). In contrast to traditional ResNet architectures, DenseResNets are feed-forward networks that allow for shortcut connections across all layers. These additional connections introduce memory in the network architecture, as typical in many modern architectures. In this work, we explore how the memory capacity in neural DDEs influences the universal approximation property. The key parameter for studying the memory capacity is the product $K τ$ of the Lipschitz constant and the delay of the DDE. In the case of non-augmented architectures, where the network width is not larger than the input and output dimensions, neural ODEs and classical feed-forward neural networks cannot have the universal approximation property. We show that if the memory capacity $Kτ$ is sufficiently small, the dynamics of the neural DDE can be approximated by a neural ODE. Consequently, non-augmented neural DDEs with a small memory capacity also lack the universal approximation property. In contrast, if the memory capacity $Kτ$ is sufficiently large, we can establish the universal approximation property of neural DDEs for continuous functions. If the neural DDE architecture is augmented, we can expand the parameter regions in which universal approximation is possible. Overall, our results show that by increasing the memory capacity $Kτ$, the infinite-dimensional phase space of DDEs with positive delay $τ>0$ is not sufficient to guarantee a direct jump transition to universal approximation, but only after a certain memory threshold, universal approximation holds.

神经ODE延迟微分通用逼近深度学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。