对比不同机器学习势模型的隐含特征,揭示其化学空间表征差异。
Comparing the latent features of universal machine-learning interatomic potentials
- 通过特征重建误差量化模型隐含特征的信息量
- 不同模型间特征重建误差大,表征方式显著不同
- 适合关注模型可解释性与化学信息压缩的研究者
近年来,通用机器学习原子间势(uMLIPs)能够在广泛化学结构和组成下以合理精度逼近基态势能面。尽管模型架构与训练数据各异,它们均能将海量化学信息压缩为描述性隐含特征。本文系统分析了不同uMLIP所学内容,通过特征重建误差定量评估其隐含特征的信息含量,并考察训练集与训练协议对趋势的影响。结果表明,uMLIPs以显著不同的方式编码化学空间,跨模型特征重建误差较大。相同架构的不同变体,其特征趋势依赖于训练数据、目标函数及训练协议。此外,微调后的模型仍保留强预训练偏差。最后,我们发现可通过逐级累积量拼接,将原子级特征压缩为全局结构级特征,每层累积量均带来关于原子环境变异性的新信息。
原文摘要 · Abstract (English)
The past few years have seen the development of ``universal'' machine-learning interatomic potentials (uMLIPs) capable of approximating the ground-state potential energy surface across a wide range of chemical structures and compositions with reasonable accuracy. While these models differ in the architecture and the dataset used, they share the ability to compress a staggering amount of chemical information into descriptive latent features. Herein, we systematically analyze what the different uMLIPs have learned by quantitatively assessing the relative information content of their latent features with feature reconstruction errors, and observing how the trends are affected by the choice of training set and training protocol. We find that uMLIPs encode the chemical space in significantly distinct ways, with substantial cross-model feature reconstruction errors. When variants of the same model architecture are considered, trends become dependent on the dataset, target, and training protocol of choice. We also observe that fine-tuning of a uMLIP retains a strong pre-training bias in the latent features. Finally, we discuss how atom-level features, which are directly output by MLIPs, can be compressed into global structure-level features via concatenation of progressive cumulants, each adding significantly new information about the variability across the atomic environments within a given system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。