arXiv:2506.18497cond-mat.mtrl-scics.LG2025-06被引 10

用预训练势能模型提取特征,实现高效精准的材料性质预测。

Leveraging neural network interatomic potentials for a foundation model of chemistry

  • 先用预训练神经网络势能模型提取固定长度特征向量,再用浅层模型预测性质。
  • 在少量数据下性能超越端到端深度网络,数据量超1万时优势明显。
  • 适合缺乏大规模数据的研究者,提升材料性质预测效率与可及性。

大规模基础模型,包括计算材料科学中的神经网络原子间势能(NIPs),已展现出巨大潜力。尽管其能加速原子级模拟,但NIPs难以直接预测电子性质,且需耦合更高尺度模型或大量模拟才能获得宏观性质。机器学习虽可实现结构-性质映射,但存在权衡:基于特征的方法泛化能力差,而深度神经网络则需大量数据和算力。为此,本文提出HackNIP,一种两阶段流程:首先从预训练的NIP基础模型中提取固定长度的特征向量(嵌入),再利用这些嵌入训练浅层机器学习模型进行下游性质预测。该研究验证了这种“破解”NIP的混合策略是否优于端到端深度网络,确定了该迁移学习方法超越直接微调NIP的数据量阈值,并识别出最优嵌入深度。HackNIP在Matbench上进行基准测试,评估其数据效率,并在包含 extit{ab initio}、实验和分子性质的多样化任务中验证。同时分析了嵌入深度对性能的影响。结果表明,该混合策略有效克服了材料科学中机器学习的权衡,旨在推动高性能预测建模的普及。

原文摘要 · Abstract (English)

Large-scale foundation models, including neural network interatomic potentials (NIPs) in computational materials science, have demonstrated significant potential. However, despite their success in accelerating atomistic simulations, NIPs face challenges in directly predicting electronic properties and often require coupling to higher-scale models or extensive simulations for macroscopic properties. Machine learning (ML) offers alternatives for structure-to-property mapping but faces trade-offs: feature-based methods often lack generalizability, while deep neural networks require significant data and computational power. To address these trade-offs, we introduce HackNIP, a two-stage pipeline that leverages pretrained NIPs. This method first extracts fixed-length feature vectors (embeddings) from NIP foundation models and then uses these embeddings to train shallow ML models for downstream structure-to-property predictions. This study investigates whether such a hybridization approach, by ``hacking" the NIP, can outperform end-to-end deep neural networks, determines the dataset size at which this transfer learning approach surpasses direct fine-tuning of the NIP, and identifies which NIP embedding depths yield the most informative features. HackNIP is benchmarked on Matbench, evaluated for data efficiency, and tested on diverse tasks including \textit{ab initio}, experimental, and molecular properties. We also analyze how embedding depth impacts performance. This work demonstrates a hybridization strategy to overcome ML trade-offs in materials science, aiming to democratize high-performance predictive modeling.

材料科学迁移学习嵌入表示预测建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。