测试80种基础模型在催化场景下的表现,发现其有高精度也有致命短板。
How accurate are foundational machine learning interatomic potentials for heterogeneous catalysis?
- 用真实催化任务测试80种机器学习势函数的零样本性能。
- 在钙钛矿空位能等任务上精度接近从头算,但磁性材料会严重失效。
- 不同任务需选不同模型,无万能通用势函数。
基础机器学习势函数(MLIPs)发展迅速,逼近从头算精度,有望实现更大尺度的模拟。然而,现有基准测试多限于有序晶体和体相材料,难以反映其在实际应用如异相催化中的表现。本文系统评估了80种不同MLIPs在异相催化典型任务中的零样本性能,涵盖合金金属、氧化物及金属-氧化物界面体系上的吸附与反应。结果表明,当前生成的先进MLIPs在预测钙钛矿氧化物空位形成能或负载纳米团簇零点能时已具高精度;但许多模型在磁性材料上灾难性失败,且结构弛豫后的能量预测误差显著高于对已优化结构的单点计算。对比低成本任务特定模型,发现若仅考虑精度,部分专用模型可媲美顶尖MLIPs。此外,无单一模型在所有任务中表现最优,用户需针对具体应用评估模型适用性。
原文摘要 · Abstract (English)
Foundational machine learning interatomic potentials (MLIPs) are being developed at a rapid pace, promising closer and closer approximation to ab initio accuracy. This unlocks the possibility to simulate much larger length and time scales. However, benchmarks for these MLIPs are usually limited to ordered, crystalline and bulk materials. Hence, reported performance does not necessarily accurately reflect MLIP performance in real applications such as heterogeneous catalysis. Here, we systematically analyze zero-shot performance of 80 different MLIPs, evaluating tasks typical for heterogeneous catalysis across a range of different data sets, including adsorption and reaction on surfaces of alloyed metals, oxides, and metal-oxide interfacial systems. We demonstrate that current-generation foundational MLIPs can already perform at high accuracy for applications such as predicting vacancy formation energies of perovskite oxides or zero-point energies of supported nanoclusters. However, limitations also exist. We find that many MLIPs catastrophically fail when applied to magnetic materials, and structure relaxation in the MLIP generally increases the energy prediction error compared to single-point evaluation of a previously optimized structure. Comparing low-cost task-specific models to foundational MLIPs, we highlight some core differences between these model approaches and show that -- if considering only accuracy -- these models can compete with the current generation of best-performing MLIPs. Furthermore, we show that no single MLIP universally performs best, requiring users to investigate MLIP suitability for their desired application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。