arXiv:2605.15179cs.LGcs.AI2026-05

通过稀疏专家路由解决多物理模型训练中的负迁移问题。

Eradicating Negative Transfer in Multi-Physics Foundation Models via Sparse Mixture-of-Experts Routing

  • 用稀疏激活的专家网络动态分配不同物理场景的计算路径。
  • 在混合流体与多孔介质数据上同时收敛,误差低至10^-6量级。
  • 适合需要统一建模多种物理机制的科学机器学习研究者。

将科学机器学习(SciML)扩展为通用基础模型受制于负迁移:同时训练不同偏微分方程(PDE)模式会引发梯度冲突、优化不稳和可塑性损失。宽带明渠流与边界主导的多孔介质流对单一密集参数路径提出矛盾的频谱与几何要求。我们提出Shodh-MoE,一种基于压缩16^3物理隐变量的稀疏激活潜变器架构,采用赫姆霍兹风格速度参数化,确保解码状态保持无散度。模型实现精确质量守恒,128^3网格上后验验证的速率散度约为2.8×10^-10(FP64)。Top-1软语义路由器将局部隐变量块分配给特定专家子网,使不同物理机制拥有专用参数路径,同时保留共享专家以维持通用对称性。在20,000步分布式预训练中,路由日志显示自动域分离:明渠验证样本仅路由至专家0,多孔介质样本仅路由至专家1。模型在两域同步收敛,隐空间验证均方误差分别为2.46×10^-5和9.76×10^-6,解码物理场均方误差为2.48×10^-6和1.76×10^-6。结果表明,稀疏专家路由是缓解通用神经算子中多物理干扰的有效架构机制。

原文摘要 · Abstract (English)

Scaling Scientific Machine Learning (SciML) toward universal foundation models is bottlenecked by negative transfer: the simultaneous co-training of disparate partial differential equation (PDE) regimes can induce gradient conflict, unstable optimization, and plasticity loss in dense neural operators. In particular, broadband open-channel fluid dynamics and boundary-dominated porous media flows impose incompatible spectral and geometric demands on a single dense parameter path. We introduce Shodh-MoE, a sparse-activated latent transformer architecture for multi-physics transport. Shodh-MoE operates on compressed 16^3 physical latents produced by a physics-informed autoencoder with an intra-tokenizer Helmholtz-style velocity parameterization, restricting decoded states to divergence-free velocity manifolds. The model guarantees exact mass conservation, achieving a physically verifiable velocity divergence of ~2.8 x 10^-10 (evaluated post-hoc in FP64) on 128^3 grids. A Top-1 soft-semantic router dynamically assigns localized latent patches to expert subnetworks, enabling specialized parameter paths for distinct physical mechanisms while preserving shared experts for universal symmetries. In a 20,000-step distributed pretraining run over mixed three-dimensional physical tensors, routing telemetry shows autonomous domain bifurcation: held-out validation tokens from the open-channel domain route exclusively to Expert 0, while porous-media tokens route exclusively to Expert 1. The model converges simultaneously across both regimes, achieving latent validation MSEs of 2.46 x 10^-5 and 9.76 x 10^-6, and decoded physical MSEs of 2.48 x 10^-6 and 1.76 x 10^-6. These results support sparse expert routing as a practical architectural mechanism for mitigating multi-physics interference in universal neural operators.

多物理建模专家路由科学机器学习神经算子

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。