arXiv:2501.16398cs.LGphysics.atom-ph2025-01

用局部原子环境差向量减少数据冗余,加速机器学习势能训练。

Data-Efficient Machine Learning Potentials via Difference Vectors Based on Local Atomic Environments

  • 基于局部原子环境的差向量编码结构差异,结合t-SNE可视化。
  • 在多种材料中实现56%数据量缩减,单次训练时间降超50%。
  • 可识别分布外数据,提升模型可靠性,适合大规模模拟研究者。

构建高效且多样的数据集对于原子模拟中高精度机器学习势能(MLPs)的发展至关重要。然而,现有方法常面临数据冗余和高计算成本问题。本文提出一种新方法——基于局部原子环境的差向量(DV-LAE),通过基于直方图的描述符编码结构差异,并利用t-SNE降维实现可视化。该方法有助于检测冗余、优化数据集,同时保持结构多样性。我们在多种材料体系中验证了其有效性,包括高压氢、铁-氢二元系统、镁氢化物和碳同素异形体。例如,在α-Fe/H系统中,维持相近的预测精度下,数据集规模减少56%,每次迭代训练时间下降超过50%。此外,通过分析高误差预测点的空间分布,可视化DV-LAE表示可有效识别分布外数据,为新结构提供可靠的置信度评估。结果表明,局部环境可视化不仅是可解释性工具,更是加速MLP开发、提升大规模原子建模数据效率的实际手段。

原文摘要 · Abstract (English)

Constructing efficient and diverse datasets is essential for the development of accurate machine learning potentials (MLPs) in atomistic simulations. However, existing approaches often suffer from data redundancy and high computational costs. Herein, we propose a new method--Difference Vectors based on Local Atomic Environments (DV-LAE)--that encodes structural differences via histogram-based descriptors and enables visual analysis through t-SNE dimensionality reduction. This approach facilitates redundancy detection and dataset optimization while preserving structural diversity. We demonstrate that DV-LAE significantly reduces dataset size and training time across various materials systems, including high-pressure hydrogen, iron-hydrogen binaries, magnesium hydrides, and carbon allotropes, with minimal compromise in prediction accuracy. For instance, in the $α$-Fe/H system, maintaining a highly similar MLP accuracy, the dataset size was reduced by 56%, and the training time per iteration dropped by over 50%. Moreover, we show how visualizing the DV-LAE representation aids in identifying out-of-distribution data by examining the spatial distribution of high-error prediction points, providing a robust reliability metric for new structures during simulations. Our results highlight the utility of local environment visualization not only as an interpretability tool but also as a practical means for accelerating MLP development and ensuring data efficiency in large-scale atomistic modeling.

机器学习势数据效率原子模拟结构表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。