arXiv:2411.15101cs.LGphysics.comp-ph2024-11被引 8

神经偏微分方程可能学的是数值误差而非真实物理。

What You See is Not What You Get: Neural Partial Differential Equations and The Illusion of Learning

  • 用数值分析方法揭示模型学习的是离散化带来的计算误差
  • 训练数据中的截断误差导致模型系统性偏差,泛化依赖偶然巧合
  • 可通过权重特征提前判断模型在新场景下的可靠性

可微编程在科学机器学习中日益流行,常将神经网络嵌入偏微分方程(NeuralPDE),被视为更可信、更具泛化能力。然而,这类模型依赖高质量的数值模拟作为“真实标签”,而这些模拟本质上是真实物理的离散近似。本文通过数值分析、实验和模型雅可比分析,发现NeuralPDE实际上学习的是数值离散化中因空间导数泰勒展开截断产生的误差。模型存在系统性偏差,其泛化能力依赖训练数据与模型间数值耗散和截断误差的偶然匹配,这在实际应用中罕见。该偏差在简单的一维方程中已显著显现,警示高维复杂真实世界问题的可信度。此外,初始条件会限制初值问题的截断误差,从而制约外推能力。最后,我们证明对模型权重进行特征值分析可提前判断其在分布外测试时的不准确性。

原文摘要 · Abstract (English)

Differentiable Programming for scientific machine learning (SciML) has recently seen considerable interest and success, as it directly embeds neural networks inside PDEs, often called as NeuralPDEs, derived from first principle physics. Therefore, there is a widespread assumption in the community that NeuralPDEs are more trustworthy and generalizable than black box models. However, like any SciML model, differentiable programming relies predominantly on high-quality PDE simulations as "ground truth" for training. However, mathematics dictates that these are only discrete numerical approximations of the true physics. Therefore, we ask: Are NeuralPDEs and differentiable programming models trained on PDE simulations as physically interpretable as we think? In this work, we rigorously attempt to answer these questions, using established ideas from numerical analysis, experiments, and analysis of model Jacobians. Our study shows that NeuralPDEs learn the artifacts in the simulation training data arising from the discretized Taylor Series truncation error of the spatial derivatives. Additionally, NeuralPDE models are systematically biased, and their generalization capability is likely enabled by a fortuitous interplay of numerical dissipation and truncation error in the training dataset and NeuralPDE, which seldom happens in practical applications. This bias manifests aggressively even in relatively accessible 1-D equations, raising concerns about the veracity of differentiable programming on complex, high-dimensional, real-world PDEs, and in dataset integrity of foundation models. Further, we observe that the initial condition constrains the truncation error in initial-value problems in PDEs, thereby exerting limitations to extrapolation. Finally, we demonstrate that an eigenanalysis of model weights can indicate a priori if the model will be inaccurate for out-of-distribution testing.

神经微分方程数值误差模型可信度可微编程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。