arXiv:2505.22208cs.LG2025-05

用3亿条半标注数据,让材料模型预训练更高效准确。

LaMM: Semi-Supervised Pre-Training of Large-Scale Materials Models

  • 引入去噪自监督学习与负载均衡算法,提升大规模预训练效率。
  • 在约3亿样本上训练单个模型,微调后速度与精度均提升。
  • 适合需要高性能材料模拟的科研人员和工业应用者。

神经网络势函数(NNPs)通过替代密度泛函理论(DFT)计算,可显著加速计算材料科学。提升其精度可通过预训练与微调实现:先在大规模数据集上预训练,再在小规模目标数据集上微调。但该方法计算成本高,主要源于DFT标注数据的开销及大规模预训练中的负载不均。为此,我们提出LaMM,一种结合改进去噪自监督学习与负载均衡算法的半监督预训练方法,有效利用约3亿条半标注样本训练单一NNP模型,显著提升了微调阶段的速度与精度。

原文摘要 · Abstract (English)

Neural network potentials (NNPs) are crucial for accelerating computational materials science by surrogating density functional theory (DFT) calculations. Improving their accuracy is possible through pre-training and fine-tuning, where an NNP model is first pre-trained on a large-scale dataset and then fine-tuned on a smaller target dataset. However, this approach is computationally expensive, mainly due to the cost of DFT-based dataset labeling and load imbalances during large-scale pre-training. To address this, we propose LaMM, a semi-supervised pre-training method incorporating improved denoising self-supervised learning and a load-balancing algorithm for efficient multi-node training. We demonstrate that our approach effectively leverages a large-scale dataset of $\sim$300 million semi-labeled samples to train a single NNP model, resulting in improved fine-tuning performance in terms of both speed and accuracy.

材料模拟半监督学习神经网络势函数预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。