arXiv:2509.21647cond-mat.mtrl-scics.LG2025-09被引 6

用大模型自动生成数据训练机器学习势,省时省力还准。

Automated Machine Learning Pipeline: Large Language Models-Assisted Automated Dataset Generation for Training Machine-Learned Interatomic Potentials

  • 大模型协助选计算代码、准备输入输出,全自动构建数据集
  • 在多晶型分子上实现1.7 meV/atom能量误差和7.0 meV/Å力误差
  • 适合想快速部署高精度分子模拟的科研人员

机器学习原子间势(MLIPs)已成为突破量子方法计算限制的强大工具,在远低于量子计算成本的前提下实现近量子精度。然而,构建可靠MLIP仍困难重重,需高质量数据集生成、原子结构预处理及精细模型训练与验证。本文提出自动化机器学习流程(AMLP),统一从数据生成到模型验证的全流程。AMLP利用大语言模型代理协助电子结构代码选择、输入准备和输出转换,其分析套件(AMLP-Analysis)基于ASE支持多种分子模拟。该流程基于MACE架构,在吖啶多晶型体系上验证:仅通过微调基础模型,即达到约1.7 meV/atom的能量均方误差和约7.0 meV/Å的力误差。拟合后的MLIP能以亚埃级精度复现DFT几何结构,并在微正则与正则系综的分子动力学模拟中表现稳定。

原文摘要 · Abstract (English)

Machine learning interatomic potentials (MLIPs) have become powerful tools to extend molecular simulations beyond the limits of quantum methods, offering near-quantum accuracy at much lower computational cost. Yet, developing reliable MLIPs remains difficult because it requires generating high-quality datasets, preprocessing atomic structures, and carefully training and validating models. In this work, we introduce an Automated Machine Learning Pipeline (AMLP) that unifies the entire workflow from dataset creation to model validation. AMLP employs large-language-model agents to assist with electronic-structure code selection, input preparation, and output conversion, while its analysis suite (AMLP-Analysis), based on ASE supports a range of molecular simulations. The pipeline is built on the MACE architecture and validated on acridine polymorphs, where, with a straightforward fine-tuning of a foundation model, mean absolute errors of ~1.7 meV/atom in energies and ~7.0 meV/Å in forces are achieved. The fitted MLIP reproduces DFT geometries with sub-Å accuracy and demonstrates stability during molecular dynamics simulations in the microcanonical and canonical ensembles.

机器学习势自动化流程大模型应用分子模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。