arXiv:2603.07977physics.chem-phcs.LG2026-03被引 2

用专家混合架构提升原子间势能模型的精度与效率

Mixture of experts architectures for machine learning interatomic potentials

  • 采用元素级稀疏路由与共享专家设计,增强模型表达能力
  • 在多个基准测试中超越所有DPA3基线,最高提升12.3%准确率
  • 专家分工符合周期表规律,可解释性强,适合材料模拟研究

机器学习原子间势能(MLIPs)可实现高精度大规模原子模拟,但高效提升其表达能力仍具挑战。本文系统研究了在DPA3框架下Mixture-of-Experts(MoE)与Mixture-of-Linear-Experts(MoLE)架构,分析路由策略与专家设计的影响。结果表明,稀疏激活结合共享专家带来显著性能提升;当存在共享专家时,非线性MoE优于MoLE,凸显非线性专家特化的必要性。元素级路由始终优于配置级路由,而全局MoE路由常引发数值不稳定性。所提出的元素级MoE模型在OMol25、OMat24和OC20M基准上持续优于所有DPA3基线。路由模式分析显示专家分工具有化学可解释性,与周期表趋势一致,表明模型有效捕捉元素特异性化学特征,实现精准原子建模。

原文摘要 · Abstract (English)

Machine Learning Interatomic Potentials (MLIPs) enable accurate large-scale atomistic simulations, yet improving their expressive capacity efficiently remains challenging. Here we systematically investigate Mixture-of-Experts (MoE) and Mixture-of-Linear-Experts (MoLE) architectures within the DPA3 framework for MLIPs and analyze the effects of routing strategies and expert designs. We show that sparse activation combined with shared experts yields substantial performance gains, and that nonlinear MoE formulations outperform MoLE when shared experts are present, underscoring the importance of nonlinear expert specialization. Furthermore, element-wise routing consistently surpasses configuration-level routing, while global MoE routing often leads to numerical instability. The resulting element-wise MoE model consistently outperforms all DPA3-based baselines across the OMol25, OMat24, and OC20M benchmarks. Analysis of routing patterns reveals chemically interpretable expert specialization aligned with periodic-table trends, indicating that the model effectively captures element-specific chemical characteristics for precise interatomic modeling.

原子势能专家混合材料模拟深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。