arXiv:2409.07610cond-mat.mtrl-scics.LG2024-09

数据越多越差?研究发现删掉氮富集结构能显著提升势函数精度。

When More Data Hurts: Optimizing Data Coverage While Mitigating Diversity Induced Underfitting in an Ultra-Fast Machine-Learned Potential

  • 用自动生成与专家数据构建训练集,对比不同多样性下的模型表现。
  • 去除氮富集结构的子集训练使预测和模拟精度大幅提升。
  • 强调针对应用需求定制数据,避免过度多样化导致欠拟合。

机器学习势函数(MLIPs)在材料建模中日益重要,但训练数据的生成优化仍是重大挑战。这是因为当分子动力学(MD)模拟遇到与训练数据差异过大的局部环境时,模型可能失效。由于无法预先确定模拟中将遇到的环境,因此需要多样且高质量的训练数据。本研究以非晶氮化硅为对象,采用超快力场(UF$^3$)评估训练数据多样性对性能的影响。通过结合专家与自主生成的数据,构建多个数据子集并训练四种力场变体。结果表明:训练数据多样性存在关键平衡——多样性不足会限制泛化能力,而过度多样化则超出模型学习容量,降低模拟精度。特别地,移除氮富集结构的子集训练出的UF$^3$变体,在预测与模拟精度上显著优于其他版本。该研究揭示了构建准确MLIP所需的精细数据设计要求,强调应根据具体应用场景定制训练数据以实现最优性能。

原文摘要 · Abstract (English)

Machine-learned interatomic potentials (MLIPs) are becoming an essential tool in materials modeling. However, optimizing the generation of training data used to parameterize the MLIPs remains a significant challenge. This is because MLIPs can fail when encountering local enviroments too different from those present in the training data. The difficulty of determining \textit{a priori} the environments that will be encountered during molecular dynamics (MD) simulation necessitates diverse, high-quality training data. This study investigates how training data diversity affects the performance of MLIPs using the Ultra-Fast Force Field (UF$^3$) to model amorphous silicon nitride. We employ expert and autonomously generated data to create the training data and fit four force-field variants to subsets of the data. Our findings reveal a critical balance in training data diversity: insufficient diversity hinders generalization, while excessive diversity can exceed the MLIP's learning capacity, reducing simulation accuracy. Specifically, we found that the UF$^3$ variant trained on a subset of the training data, in which nitrogen-rich structures were removed, offered vastly better prediction and simulation accuracy than any other variant. By comparing these UF$^3$ variants, we highlight the nuanced requirements for creating accurate MLIPs, emphasizing the importance of application-specific training data to achieve optimal performance in modeling complex material behaviors.

机器学习势数据多样性材料模拟非晶材料

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。