用正交多项式约束权重变化,提升连续时间神经网络的稳定性和效率
Weight-Parameterization in Continuous Time Deep Neural Networks for Surrogate Modeling
- 用多项式基函数限制权重随时间变化的轨迹
- 勒让德基比单项式基更稳定,计算成本更低,精度相当或更高
- 适合需要高效建模复杂物理系统的研究人员
连续时间深度学习模型(如神经微分方程)为复杂物理系统的代理建模提供了有前景的框架。训练中的核心挑战在于学习表达性强且稳定的时变权重,尤其在计算资源受限条件下。本文研究了将权重的时间演化限制在由多项式基函数张成的低维子空间中的参数化策略。在神经ODE和残差网络(ResNet)架构中,对比了单项式与勒让德多项式基,在离散后优化和优化后离散两种训练范式下的表现。在三个高维基准问题上的实验结果表明,勒让德参数化能带来更稳定的训练动态,降低计算开销,并实现与单项式参数化及无约束权重模型相当或更优的精度。这些发现揭示了基函数选择在时变权重参数化中的关键作用,证明使用正交多项式基可在模型表达力与训练效率间取得良好平衡。
原文摘要 · Abstract (English)
Continuous-time deep learning models, such as neural ordinary differential equations (ODEs), offer a promising framework for surrogate modeling of complex physical systems. A central challenge in training these models lies in learning expressive yet stable time-varying weights, particularly under computational constraints. This work investigates weight parameterization strategies that constrain the temporal evolution of weights to a low-dimensional subspace spanned by polynomial basis functions. We evaluate both monomial and Legendre polynomial bases within neural ODE and residual network (ResNet) architectures under discretize-then-optimize and optimize-then-discretize training paradigms. Experimental results across three high-dimensional benchmark problems show that Legendre parameterizations yield more stable training dynamics, reduce computational cost, and achieve accuracy comparable to or better than both monomial parameterizations and unconstrained weight models. These findings elucidate the role of basis choice in time-dependent weight parameterization and demonstrate that using orthogonal polynomial bases offers a favorable tradeoff between model expressivity and training efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。