证明了神经PDE在伴随梯度下降下的全局收敛性,为科学机器学习提供理论支撑。
Global Convergence of Adjoint-Optimized Neural PDEs
- 用伴随方程高效计算梯度,实现神经PDE参数优化
- 在隐层无限宽与训练时间无穷长下,解逼近目标数据
- 适用于需补全物理机制的科学建模,如流体、热传导
许多工程与科学领域正尝试用神经网络建模偏微分方程(PDE)中的未知项,需从观测数据中反推神经网络项以近似缺失或未解析的物理过程。由此产生的神经网络PDE模型依赖于网络参数,可通过梯度下降法在满足真实数据的条件下进行校准,其中梯度通过求解伴随PDE实现高效计算。本文研究了在隐藏单元数和训练时间趋于无穷时,伴随梯度下降方法在训练神经网络PDE模型中的收敛性。针对一类包含神经网络源项的非线性抛物型PDE,证明了训练后的神经网络PDE解可收敛至目标数据(即全局极小值)。该全局收敛性证明面临独特数学挑战:一是无限宽隐层极限下神经网络训练动态涉及非局部核算子,其特征值无谱间隙;二是极限PDE系统的非线性导致即使在无限宽度下,神经网络函数优化问题仍为非凸(不同于典型神经网络在大神经元极限下趋于凸的情况)。理论结果通过数值实验验证。
原文摘要 · Abstract (English)
Many engineering and scientific fields have recently become interested in modeling terms in partial differential equations (PDEs) with neural networks, which requires solving the inverse problem of learning neural network terms from observed data in order to approximate missing or unresolved physics in the PDE model. The resulting neural-network PDE model, being a function of the neural network parameters, can be calibrated to the available ground truth data by optimizing over the PDE using gradient descent, where the gradient is evaluated in a computationally efficient manner by solving an adjoint PDE. These neural PDE models have emerged as an important research area in scientific machine learning. In this paper, we study the convergence of the adjoint gradient descent optimization method for training neural PDE models in the limit where both the number of hidden units and the training time tend to infinity. Specifically, for a general class of nonlinear parabolic PDEs with a neural network embedded in the source term, we prove convergence of the trained neural-network PDE solution to the target data (i.e., a global minimizer). The global convergence proof poses a unique mathematical challenge that is not encountered in finite-dimensional neural network convergence analyses due to (i) the neural network training dynamics involving a non-local neural network kernel operator in the infinite-width hidden layer limit where the kernel lacks a spectral gap for its eigenvalues and (ii) the nonlinearity of the limit PDE system, which leads to a non-convex optimization problem in the neural network function even in the infinite-width hidden layer limit (unlike in typical neural network training cases where the optimization problem becomes convex in the large neuron limit). The theoretical results are illustrated and empirically validated by numerical studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。