通过自适应变换简化复杂微分方程解算子,提升通用模型性能。
AOT-POT: Adaptive Operator Transformation for Large-Scale PDE Pre-training

- 设计自适应算子变换机制,动态重构不同PDE解的结构形式。
- 在12个基准上实现领先性能,平均降低40.9%相对L2误差。
- 适合需要跨类型微分方程建模的研究者与工程应用者。
在多样化的偏微分方程(PDE)数据集上预训练神经算子,已成为构建科学机器学习中通用代理模型的有前景方向。然而,PDE解算子固有的复杂性和结构多样性使多PDE预训练面临根本挑战。现有方法主要通过增加模型容量来应对,而保持目标解算子不变。受经典数值分析启发,我们提出将复杂多样的解算子转化为更简单、对齐性更强的形式,以更易联合建模。由于最优变换随PDE类型变化,必须具备自适应和输入依赖性,使单一神经算子可逼近整个算子族。我们提出AOT-POT(自适应算子变换预训练算子变压器),通过扩展隐藏表示为多并行流,自适应聚合与重分配各子层前后信息,并使用Sinkhorn投影双随机矩阵混合流,实现稳定训练。这些机制共同将多样解算子重塑为统一形式,由单一架构有效建模。实证表明,AOT-POT在12个PDE基准上表现最优,仅增加3%参数,相对L2误差最高降低77.6%(平均降低40.9%)。微调后,对域内PDE误差再降92%,对域外(预训练未见类型)误差降低89%,证明自适应算子变换是超越单纯扩大模型容量的有效补充方向。
原文摘要 · Abstract (English)
Pre-training neural operators on diverse partial differential equation (PDE) datasets has emerged as a promising direction for building general-purpose surrogate models in scientific machine learning. However, the inherent complexity and structural diversity of PDE solution operators make multi-PDE pre-training fundamentally challenging. Existing methods mainly address this by increasing model capacity, while leaving the target solution operators unchanged. Inspired by classical numerical analysis, we instead propose to transform complex and diverse solution operators into simpler, better-aligned forms that are easier to model jointly. Since the optimal transformation varies across PDE types, it must be adaptive and input-dependent, allowing a single neural operator to approximate an entire family of operators. We instantiate this idea as AOT-POT (adaptive operator-transformation for pre-training operator transformer), which expands hidden representations into multiple parallel streams, adaptively aggregates and redistributes them before and after each sub-layer, and mixes streams through Sinkhorn-projected doubly stochastic matrices for stable training. These mechanisms together reshape diverse solution operators into a unified form that can be effectively modeled by a single architecture. Empirically, AOT-POT achieves state-of-the-art performance on 12 PDE benchmarks with only 3\% additional parameters, reducing relative L2 error by up to 77.6\% (40.9\% on average). Fine-tuning AOT-POT further reduces L2 error by up to 92\% on in-domain PDEs and 89\% on out-of-domain PDEs (unseen types during pre-training), demonstrating that adaptive operator transformation is an effective and complementary direction for advancing PDE foundation models beyond simply scaling model capacity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。