用分子力大小指导采样,让模型更准更省数据。
Gradient-Guided Furthest Point Sampling for Robust Training Set Selection
- 基于分子力大小优化采样点分布,避免传统方法漏掉关键构型。
- 在二维测试和MD17数据上,训练量减半仍保持精度,误差显著降低。
- 适合追求高效与鲁棒性的化学机器学习研究者使用。
训练集采样方法可提升化学相关机器学习任务的模型性能并降低数据成本。本文提出梯度引导最远点采样(GGFPS),作为最远点采样(FPS)的简单扩展,利用分子力范数引导分子构型空间的高效采样。通过一个简化系统(Styblinski-Tang函数)以及来自MD17数据集的分子动力学轨迹,验证了该方法的有效性。结果表明,相较于FPS和均匀随机采样(URS),以及已有的监督式FPS变体PCov-FPS和PCov-CUR,GGFPS展现出更高的数据效率和更强的模型鲁棒性。对MD17数据的分布分析显示,传统FPS会系统性低估平衡构型,导致松弛结构预测误差较大。而GGFPS有效纠正这一偏差,实现:(i) 在二维Styblinski-Tang系统中,训练成本减半而不损失预测精度;(ii) 在MD17中系统性降低平衡与应变结构的预测误差;(iii) 显著减小所有构型空间下的预测误差方差。这些结果表明,梯度感知采样方法具有成为高效训练集选择工具的巨大潜力,而直接使用传统FPS可能导致训练不平衡和预测不一致。
原文摘要 · Abstract (English)
Training set sampling methods are used to improve model performance and lower data costs in machine learning problems relevant to chemistry. We introduce Gradient Guided Furthest Point Sampling (GGFPS), a simple extension of Furthest Point Sampling (FPS) that leverages molecular force norms to guide efficient sampling of configurational spaces of molecules. Numerical evidence is presented for a toy system (the Styblinski-Tang function) as well as for molecular dynamics trajectories from the MD17 dataset. Our numerical results indicate superior data efficiency and model robustness when using GGFPS compared to FPS and uniform random sampling (URS), as well as established supervised FPS-style selectors, PCov-FPS and PCov-CUR. Distribution analysis of the MD17 data suggests that FPS systematically under-samples equilibrium geometries, resulting in large test errors for relaxed structures. GGFPS cures this artifact and (i) enables up to twofold reductions in training cost without sacrificing predictive accuracy compared to FPS in the 2-dimensional Styblinski-Tang system, (ii) systematically lowers prediction errors for equilibrium as well as strained structures in MD17, and (iii) systematically decreases prediction error variances across all of the MD17 configuration spaces. These results suggest that gradient-aware sampling methods hold great promise as effective training set selection tools, and that naive use of FPS may result in imbalanced training and inconsistent prediction outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。