无需梯度信息,用模拟数据训练神经网络解高维非线性偏微分方程
A Zeroth-Order Deep Learning Method for Fully Nonlinear Parabolic Partial Differential Equations with Unknown Coefficients
- 通过扰动蒙特卡洛轨迹构造零阶导数估计器,仅用函数值生成梯度与海森矩阵目标
- 在高维下实现稳定求解,理论证明误差由离散化、逼近、统计与零阶偏差共同构成
- 适合黑箱环境中的数据驱动偏微分方程求解,尤其适用于强化学习等复杂系统
高维未知系数的非线性抛物型偏微分方程广泛存在于科学机器学习中,如连续时间强化学习,但其数据驱动求解仍具挑战。现有深度学习方法依赖重复自动微分,易引发不稳定并放大导数误差;基于随机表示的概率方法则需明确数据生成机制,无法应用于黑箱场景。本文引入两类模拟器作为数据生成机制,采用“先表示后学习”策略,在仅可通过模拟和点值评估访问底层微分算子条件下学习解及其导数。导数表示基于扰动蒙特卡洛轨迹的零阶导数(ZOD)估计器,该完全无模型方法仅通过函数值生成梯度与海森网络目标。我们提供了统计学习分析,包含ZOD的偏差-方差权衡。假设算子具有标准压缩性质,建立了非渐近误差界,将总误差分解为离散化误差、逼近误差、统计误差与ZOD偏差。关键地,推导了加权Sobolev空间中学习表示的样本复杂度,刻画至二阶导数的误差。数值实验表明该方法在中等与高维场景下表现优异。
原文摘要 · Abstract (English)
High-dimensional partial differential equations (PDEs) with unknown coefficients arise widely in scientific machine learning, including continuous-time reinforcement learning, yet solving them efficiently in a data-driven way remains challenging. Existing deep learning solvers often rely on repeated automatic differentiation to evaluate differential operators, which can cause instability and amplify derivative errors in high dimensions, while probabilistic methods based on stochastic representations require explicit knowledge of the data-generating dynamics and therefore do not apply to black-box environments. We introduce two types of simulators as data-generating mechanisms, and take a ``representing-then-learning" approach that learns the solutions and their derivatives under settings where the underlying PDE operators are accessible only through simulations and pointwise evaluations. Our representation of derivatives relies on the zeroth-order derivative (ZOD) estimators derived from perturbed Monte Carlo trajectories. This fully model-free approach generates targets for the gradient and Hessian networks using only function evaluations. We provide a statistical learning analysis of the proposed approach, including a bias--variance tradeoff for ZODs. Assuming a standard contraction property of the underlying operator, we establish a non-asymptotic error bound that decomposes the total error into discretization error, approximation error, statistical error, and ZOD bias. Crucially, we derive the sample complexity of the learned representations in (weighted) Sobolev space, characterizing the error up to second-order derivatives. Numerical experiments illustrate the competitive performance of the method in moderate and high dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。