用行列式点过程精选原子构型,提升机器学习势能模型训练效率
Data Curation for Machine Learning Interatomic Potentials by Determinantal Point Processes
- 基于分子描述符核函数的行列式点过程,选择信息量大的原子构型
- 在铪氧化物数据上,训练集精度与鲁棒性优于现有方法
- 适合异构数据或在线主动学习场景,可迭代优化数据集
机器学习势能模型的开发面临生成和标注高质量训练数据的计算瓶颈。本文提出将行列式点过程(DPPs)应用于选择需用高成本量子力学方法标注能量和力的原子构型子集。通过铪氧化物数据实验表明,利用分子描述符核函数的DPPs可构建紧凑且多样化的训练集,显著提升机器学习对分子体系表示的准确性和鲁棒性。本工作揭示了在异构或多元数据下,采用DPPs进行无监督数据筛选的潜力,也可用于分子动力学模拟中的在线主动学习,实现数据的迭代增强。
原文摘要 · Abstract (English)
The development of machine learning interatomic potentials faces a critical computational bottleneck with the generation and labeling of useful training datasets. We present a novel application of determinantal point processes (DPPs) to the task of selecting informative subsets of atomic configurations to label with reference energies and forces from costly quantum mechanical methods. Through experiments with hafnium oxide data, we show that DPPs are competitive with existing approaches to constructing compact but diverse training sets by utilizing kernels of molecular descriptors, leading to improved accuracy and robustness in machine learning representations of molecular systems. Our work identifies promising directions to employ DPPs for unsupervised training data curation with heterogeneous or multimodal data, or in online active learning schemes for iterative data augmentation during molecular dynamics simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。