研究机器学习势函数中不确定性量化的影响因素,提出改进方法提升对新原子构型的识别能力。
Model Accuracy and Data Heterogeneity Shape Uncertainty Quantification in Machine Learning Interatomic Potentials
- 用聚类增强的局部D-最优法划分构型空间,提升不确定性估计精度
- 数据异质性会降低不确定性预测效果,高精度模型可增强误差相关性
- 适用于需要可靠主动学习的机器学习势函数开发场景
机器学习原子间势函数(MLIPs)实现精准原子建模,但可靠的不确定性量化(UQ)仍难以实现。本研究在原子簇展开框架下考察了集成学习与D-最优性两种UQ策略。结果表明,模型精度越高,预测不确定性与实际误差的相关性越强,且新颖性检测能力提升,其中D-最优性给出更保守的估计。两种方法在同质训练集上均表现校准良好,但在异质数据集上低估误差且新颖性敏感度下降。为此,我们提出聚类增强的局部D-最优性方法,在训练阶段将构型空间划分为多个簇,并在每个簇内应用D-最优性。该方法显著提升了在异质数据集中对新型原子环境的检测能力。研究揭示了模型保真度与数据异质性在UQ性能中的作用,为MLIP开发提供了稳健的主动学习与自适应采样路径。
原文摘要 · Abstract (English)
Machine learning interatomic potentials (MLIPs) enable accurate atomistic modelling, but reliable uncertainty quantification (UQ) remains elusive. In this study, we investigate two UQ strategies, ensemble learning and D-optimality, within the atomic cluster expansion framework. It is revealed that higher model accuracy strengthens the correlation between predicted uncertainties and actual errors and improves novelty detection, with D-optimality yielding more conservative estimates. Both methods deliver well calibrated uncertainties on homogeneous training sets, yet they underpredict errors and exhibit reduced novelty sensitivity on heterogeneous datasets. To address this limitation, we introduce clustering-enhanced local D-optimality, which partitions configuration space into clusters during training and applies D-optimality within each cluster. This approach substantially improves the detection of novel atomic environments in heterogeneous datasets. Our findings clarify the roles of model fidelity and data heterogeneity in UQ performance and provide a practical route to robust active learning and adaptive sampling strategies for MLIP development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。