纠正材料数据库中不一致的U修正,提升机器学习势能模型的准确性
Better without U: Impact of Selective Hubbard U Correction on Foundational MLIPs
- 对含U修正的原子结构统一调整能量,消除不同势能面间的不兼容性
- 修正后模型吸附氧的能量误差显著降低,避免金属与氧之间的虚假排斥
- 适用于催化、氧化等涉及过渡金属的研究,尤其对OMAT类模型改进明显
基础机器学习势能(fMLIPs)训练依赖于通过从头计算方法生成的能量和力数据。我们发现,基于MPtrj、Alexandria和OMat24等大规模数据集训练的fMLIPs,会继承材料项目(Materials Project)选择性应用Hubbard U修正所引入的不一致性:仅当体系含O或F原子时才对特定过渡金属施加+U修正。这种不一致导致两个不相容的势能面(PES)并存——较低能的GGA表面与较高能的GGA+U表面。当模型同时在两类数据上训练时,会在二者之间进行插值,引发系统性欠键合,甚至在含U金属与含氧/氟物种间产生虚假排斥。如MACE-OMAT和-MPA模型便出现此类问题,限制其在催化与氧化研究中的应用。我们发现病理严重程度与含氧密度相关,解释了为何OMAT训练模型受影响最重,并提示未来扩展数据集若包含低氧配置(如多元素或缺陷体系),问题可能加剧。本文提出一种简单方案:按每个含U原子进行能量偏移,使相同结构的PBE+U与PBE能量对齐,获得更平滑的势能面,优于以往以相图精度为目标的修正方法。使用该修正后训练的模型,在含U元素表面吸附氧的均方误差显著减小。由于完全忽略+U的数据集(如MatPES、MP-ALOE)可避免此问题,建议未来fMLIP数据集应剔除+U;对于已有数据,本方法提供低成本后处理优化。
原文摘要 · Abstract (English)
The training of foundational machine learning interatomic potentials (fMLIPs) relies on diverse databases with energies and forces calculated using ab initio methods. We show that fMLIPs trained on large datasets such as MPtrj, Alexandria, and OMat24 encode inconsistencies from the Materials Project's selective use of the Hubbard U correction, which is applied to certain transition metals only if O or F atoms are present in the simulation cell. This inconsistent use of +U creates two incompatible potential-energy surfaces (PES): a lower-energy GGA surface and a higher-energy GGA+U one. When trained on both, MLIPs interpolate between them, leading to systematic underbinding, or even spurious repulsion, between U-corrected metals and oxygen- or fluorine-containing species. Models such as MACE-OMAT and -MPA exhibit repulsion between U-corrected metals and their oxides, limiting their value for studying catalysis and oxidation. We link the severity of this pathology to the oxygen number density in U-corrected training configurations. This explains why OMAT-trained models are most affected and suggests the issue might worsen as expanding future datasets increasingly include configurations with low oxygen content, such as those generated through combinatorial exploration of multi-element or defect-containing systems. Our simple per-U-corrected-atom shift aligns PBE+U and PBE energies for identical structures, yielding a smoother PES compared to existing correction schemes, which target phase diagram accuracy. As a result, models trained on datasets with our shift applied exhibit smaller mean absolute errors for the adsorption energies of oxygen on U-corrected elemental slabs. Since datasets omitting +U entirely (e.g. MatPES, MP-ALOE) avoid these pathologies, we recommend excluding +U in future fMLIP datasets. For existing datasets, our post-hoc correction provides a low-cost improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。