用多分辨率采样降低神经算子训练成本,提升效率。
Multi-Level Monte Carlo Training of Neural Operators
- 通过多级蒙特卡洛框架,用低分辨率数据校正梯度,减少高分辨率计算量。
- 实验表明,相同精度下训练速度提升显著,存在准确率与耗时的权衡曲线。
- 适用于各类支持多分辨率输入的神经算子模型,尤其适合大规模高分辨问题。
算子学习是快速发展的领域,旨在使用神经算子近似与偏微分方程(PDE)相关的非线性算子。这些方法依赖于函数的离散化输入输出,通常在大规模高分辨率问题上训练成本高昂。为此,我们提出一种多级蒙特卡洛(MLMC)方法,利用函数离散化的多分辨率层次结构来训练神经算子。该框架通过使用少量高分辨率数据的梯度修正,显著降低训练计算成本,同时保持高精度。所提出的MLMC训练流程可应用于任何接受多分辨率数据的模型架构。在一系列先进模型和测试案例上的数值实验表明,相比传统单分辨率训练方法,该方法在计算效率上表现更优,并揭示了准确率与计算时间之间的帕累托曲线,其特性取决于每分辨率的采样数量。
原文摘要 · Abstract (English)
Operator learning is a rapidly growing field that aims to approximate nonlinear operators related to partial differential equations (PDEs) using neural operators. These rely on discretization of input and output functions and are, usually, expensive to train for large-scale problems at high-resolution. Motivated by this, we present a Multi-Level Monte Carlo (MLMC) approach to train neural operators by leveraging a hierarchy of resolutions of function discretization. Our framework relies on using gradient corrections from fewer samples of fine-resolution data to decrease the computational cost of training while maintaining a high level accuracy. The proposed MLMC training procedure can be applied to any architecture accepting multi-resolution data. Our numerical experiments on a range of state-of-the-art models and test-cases demonstrate improved computational efficiency compared to traditional single-resolution training approaches, and highlight the existence of a Pareto curve between accuracy and computational time, related to the number of samples per resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。