突破百亿参数通用势能模型训练瓶颈,实现秒级训练。
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials

- 构建基于不变架构的百亿参数专家混合模型MatRIS-MoE
- 研发首个高维分布式训练框架Janus,支持爆破式计算优化
- 在超算上实现90%并行效率,训练时间从周级缩至小时级
通用机器学习原子间势能(uMLIPs)在涵盖整个元素周期表的无机材料与有机分子海量数据上预训练,是量子精度物理模拟的基础模型。然而,uMLIP训练需二阶导数,缺乏相应并行训练框架;扩展至百亿参数规模时,计算与通信开销呈指数增长,训练极为困难。本文提出基于不变架构的百亿参数专家混合模型MatRIS-MoE,以及首个面向uMLIP的高维分布式训练框架Janus,具备硬件感知优化能力。部署于两台埃级超算,代码在单精度下达到1.2/1.0 EFLOPS峰值性能(理论峰值的24%/35.5%),并行效率超过90%,将百亿参数uMLIP训练时间从数周压缩至数小时。该工作确立了埃级计算下AI for Science基础模型的新标杆,为快速科学发现提供关键基础设施。
原文摘要 · Abstract (English)
Universal Machine Learning Interatomic Potentials (uMLIPs), pre-trained on massively diverse datasets encompassing inorganic materials and organic molecules across the entire periodic table, serve as foundational models for quantum-accurate physical simulations. However, uMLIP training requires second-order derivatives, which lack corresponding parallel training frameworks; moreover, scaling to the billion-parameter regime causes explosive growth in computation and communication overhead, making its training a tremendous challenge. We introduce MatRIS-MoE, a billion-parameter Mixture-of-Experts model built upon invariant architecture, and {Janus}, a pioneering high-dimensional distributed training framework for uMLIPs with hardware-aware optimizations. Deployed across two Exascale supercomputers, our code attains a peak performance of 1.2/1.0 EFLOPS (24\%/{35.5\%} of theoretical peak) in single precision at over 90\% parallel efficiency, compressing the training of billion-parameter uMLIPs from weeks to hours. This work establishes a new high-water mark for AI-for-Science (AI4S) foundation models at Exascale and provides essential infrastructure for rapid scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。