不依赖过参数化,实现两层物理信息神经网络的高效优化与泛化分析
Optimization and generalization analysis for two-layer physics-informed neural networks without over-parametrization
- 在非过参数化假设下分析SGD训练两层PINN的优化行为
- 网络宽度超过阈值后,训练与期望损失均低于O(ε)
- 适用于追求理论严谨性与计算效率的科学计算研究者
本文研究随机梯度下降(SGD)在求解带有物理信息的神经网络(PINNs)最小二乘回归问题时的行为。以往相关工作基于过参数化假设,其收敛性可能要求网络宽度随训练样本数大幅增加,导致计算成本过高,与实际实验脱节。本文对两层PINN的SGD进行新的优化与泛化分析,通过关于目标函数的合理假设避免过参数化。给定任意ε>0,若网络宽度超过仅依赖于ε和问题本身的阈值,则训练损失与期望损失将降至O(ε)以下。
原文摘要 · Abstract (English)
This work focuses on the behavior of stochastic gradient descent (SGD) in solving least-squares regression with physics-informed neural networks (PINNs). Past work on this topic has been based on the over-parameterization regime, whose convergence may require the network width to increase vastly with the number of training samples. So, the theory derived from over-parameterization may incur prohibitive computational costs and is far from practical experiments. We perform new optimization and generalization analysis for SGD in training two-layer PINNs, making certain assumptions about the target function to avoid over-parameterization. Given $ε>0$, we show that if the network width exceeds a threshold that depends only on $ε$ and the problem, then the training loss and expected loss will decrease below $O(ε)$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。