用分层专家网络提升大规模微分方程求解的泛化能力
NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-Training
- 采用嵌套式混合专家架构,分别捕捉全局与局部依赖关系
- 在12个不同来源PDE数据集上预训练,下游任务表现优异
- 适合需要高泛化能力的科学计算与物理模拟场景
神经算子已成为求解偏微分方程(PDE)的高效范式,克服了传统数值方法的局限,显著提升计算效率。然而,由于PDE系统的多样性和复杂性,现有神经算子通常依赖单一网络结构,难以充分捕捉异质特征与复杂系统依赖关系,限制了基于神经算子的大规模PDE预训练。为此,我们提出一种基于嵌套混合专家(Nested MoE)框架的大规模PDE预训练神经算子。具体而言,图像级MoE用于捕获全局依赖,而令牌级子MoE专注于局部依赖。模型可针对输入动态激活最适专家网络,从而增强泛化与迁移能力。我们在来自多个来源的12个PDE数据集上进行大规模预训练,并成功将其迁移到下游任务。大量实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Neural operators have emerged as an efficient paradigm for solving PDEs, overcoming the limitations of traditional numerical methods and significantly improving computational efficiency. However, due to the diversity and complexity of PDE systems, existing neural operators typically rely on a single network architecture, which limits their capacity to fully capture heterogeneous features and complex system dependencies. This constraint poses a bottleneck for large-scale PDE pre-training based on neural operators. To address these challenges, we propose a large-scale PDE pre-trained neural operator based on a nested Mixture-of-Experts (MoE) framework. In particular, the image-level MoE is designed to capture global dependencies, while the token-level Sub-MoE focuses on local dependencies. Our model can selectively activate the most suitable expert networks for a given input, thereby enhancing generalization and transferability. We conduct large-scale pre-training on twelve PDE datasets from diverse sources and successfully transfer the model to downstream tasks. Extensive experiments demonstrate the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。