自动化构建高质量神经演化势训练数据集的工具包
NepTrain and NepTrainKit: Automated Active Learning and Visualization Toolkit for Neuroevolution Potentials
- 通过键长过滤法自动剔除非物理结构,保证数据质量
- 集成异常值检测与最远点采样,提升数据集代表性
- 适合材料模拟与机器学习势研发人员快速上手
神经演化势(NEP)具有优异的计算效率,已成功应用于材料科学。构建高质量训练数据集是开发精确NEP模型的关键,但其准备与筛选过程耗时、费力且资源消耗大,成为广泛应用的瓶颈。本文提出NepTrain和NepTrainKit,分别作为开源Python工具包与图形化界面软件,用于初始化与管理训练数据集,实现高质数据集自动生成与NEP模型训练自动化。NepTrain采用键长过滤方法有效识别并移除分子动力学轨迹中的非物理结构,确保数据质量;NepTrainKit提供数据编辑、可视化与交互探索功能,集成异常值检测、最远点采样、非物理结构识别及构型类型选择等核心特性。以CsPbI₃为例,完整演示了使用NepTrain训练NEP模型的流程,并通过材料性质预测验证模型性能。该工具包将显著提升机器学习原子间势研究者的效率。
原文摘要 · Abstract (English)
As a machine-learned potential, the neuroevolution potential (NEP) method features exceptional computational efficiency and has been successfully applied in materials science. Constructing high-quality training datasets is crucial for developing accurate NEP models. However, the preparation and screening of NEP training datasets remain a bottleneck for broader applications due to their time-consuming, labor-intensive, and resource-intensive nature. In this work, we have developed NepTrain and NepTrainKit, which are dedicated to initializing and managing training datasets to generate high-quality training sets while automating NEP model training. NepTrain is an open-source Python package that features a bond length filtering method to effectively identify and remove non-physical structures from molecular dynamics trajectories, thereby ensuring high-quality training datasets. NepTrainKit is a graphical user interface (GUI) software designed specifically for NEP training datasets, providing functionalities for data editing, visualization, and interactive exploration. It integrates key features such as outlier identification, farthest-point sampling, non-physical structure detection, and configuration type selection. The combination of these tools enables users to process datasets more efficiently and conveniently. Using $\rm CsPbI_3$ as a case study, we demonstrate the complete workflow for training NEP models with NepTrain and further validate the models through materials property predictions. We believe this toolkit will greatly benefit researchers working with machine learning interatomic potentials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。