比较神经与多项式代理模型在偏微分方程中的表现,指导选型。
Performance of Neural and Polynomial Operator Surrogates
- 对比神经与多项式代理模型的建模思路
- 光滑输入下多项式方法数据效率更高,粗糙输入下傅里叶神经算子更快收敛
- 导数信息训练提升低数据场景性能,适合有梯度信息的场景
针对参数化偏微分方程中解映射的代理算子构建问题,当正向模型计算代价高时,本文系统比较了神经算子代理(包括基于$L^2_μ$和$H^1_μ$目标的降基神经算子、傅里叶神经算子)与多项式代理方法(降基稀疏网格代理、降基张量列车代理)。所有方法在含代数衰减谱系数的输入场(衰减速率$s$变化)下,于线性参数扩散问题与非线性参数超弹性问题上进行评估。为实现公平比较,通过调整超参数生成代理模型集合,分析成本与近似精度的帕累托前沿,分解成本为数据生成、设置与评估三部分。结果表明:无方法始终最优;光滑输入($s \geq 2$)下,多项式代理具显著数据效率优势,稀疏网格代理收敛率符合理论预测;粗糙输入($s \leq 1$)下,傅里叶神经算子收敛最快。导数信息训练显著优于标准$L^2_μ$训练,在低数据且梯度可得时具竞争力。研究强调应根据问题正则性、精度需求与计算约束匹配代理方法。
原文摘要 · Abstract (English)
We consider the problem of constructing surrogate operators for parameter-to-solution maps arising from parametric partial differential equations, where repeated forward model evaluations are computationally expensive. We present a systematic empirical comparison of neural operator surrogates, including a reduced-basis neural operator trained with $L^2_μ$ and $H^1_μ$ objectives and the Fourier neural operator, against polynomial surrogate methods, specifically a reduced-basis sparse-grid surrogate and a reduced-basis tensor-train surrogate. All methods are evaluated on a linear parametric diffusion problem and a nonlinear parametric hyperelasticity problem, using input fields with algebraically decaying spectral coefficients at varying rates of decay $s$. To enable fair comparisons, we analyze ensembles of surrogate models generated by varying hyperparameters and compare the resulting Pareto frontiers of cost versus approximation accuracy, decomposing cost into contributions from data generation, setup, and evaluation. Our results show that no single method is universally superior. Polynomial surrogates achieve substantially better data efficiency for smooth input fields ($s \geq 2$), with convergence rates for the sparse-grid surrogate in agreement with theoretical predictions. For rough inputs ($s \leq 1$), the Fourier neural operator displays the fastest convergence rates. Derivative-informed training consistently improves data efficiency over standard $L^2_μ$ training, providing a competitive alternative for rough inputs in the low-data regime when Jacobian information is available at reasonable cost. These findings highlight the importance of matching the surrogate methodology to the regularity of the problem as well as accuracy demands and computational constraints of the application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。