arXiv:2609.07814cs.LGphysics.comp-ph2026-09

提出新型专家混合模型,解决多物理阶段偏微分方程求解中的训练不稳问题。

Latent-MoE: Domain-Aware Mixture-of-Experts for PDEs with Multi-Regime Physics

论文配图:Latent-MoE: Domain-Aware Mixture-of-Experts for PDEs with Multi-Regime Physics
图 1 · 摘自论文原文
  • 设计域感知专家路由,结合共享主干网络实现局部学习与能力流动
  • 在多阶段时变物理问题上性能提升超一个数量级,梯度冲突显著降低
  • 适合复杂物理场景建模,尤其适用于跨域物理变化剧烈的系统

物理信息神经网络(PINNs)在物理规律随域变化的偏微分方程求解中表现不佳。我们发现其根源在于标准坐标网络的神经正切核(NTK)具有平移非不变性,导致大坐标值的训练点对远处预测产生过度影响,引发长程耦合与梯度冲突。我们从理论上和实证上证明:采用中心化、紧支撑路由的混合专家(MoE)架构可生成均匀带状的NTK,其核回归权重随距离指数衰减,实现学习局域化。基于此,我们提出Latent-MoE,将域感知的MoE模块嵌入共享主干网络。与FB-PINNs或X-PINNs等刚性划分域与参数的方法不同,Latent-MoE在保持域感知路由局域优势的同时,通过共享主干允许容量跨区域流动。在均质物理基准测试中表现与现有方法相当;在具有多阶段时变物理的基准测试中,全局模型与刚性域分解均陷入伪解,而该方法性能提升超过一个数量级,训练过程梯度冲突明显减少。

原文摘要 · Abstract (English)

Physics-informed neural networks (PINNs) struggle on PDEs whose governing physics varies across the domain. We trace this to a structural property of standard coordinate networks: their neural tangent kernel (NTK) is translation-variant and lets training points of large coordinate magnitude disproportionately influence predictions elsewhere, producing long-range coupling and gradient conflict during training. We show analytically and empirically that mixture-of-experts (MoE) architectures with centered, compact-support routers yield a uniformly banded NTK whose kernel-regression weights decay exponentially with distance, localizing the learning. Building on this, we propose \emph{Latent-MoE}, which interleaves domain-aware MoE blocks within a shared backbone. Unlike FB-PINNs or X-PINNs, which rigidly partition both the domain and the parameters so that the parameters on different subdomains are updated independently, Latent-MoE is designed to preserve the localization benefit of domain-aware routing while allowing capacity to flow across regions through the shared backbone. On standard homogeneous-physics benchmarks Latent-MoE is competitive with established baselines; on benchmarks with multi-stage time-variable physics, where global models and rigid domain decompositions both fall into spurious solutions, it improves over them by more than an order of magnitude, with markedly reduced gradient conflict during training.

偏微分方程专家混合物理信息网络域感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。