arXiv:2510.22938cond-mat.mtrl-scics.LG2025-10被引 8

构建1350万条数据集,提升催化反应的自旋敏感性模拟精度

AQCat25: Unlocking spin-aware, high-fidelity machine learning potentials for heterogeneous catalysis

  • 引入包含1350万条DFT计算的AQCat25数据集,聚焦自旋极化系统
  • 联合训练策略在保持通用性前提下,显著提升新数据上的预测准确率
  • 通过特征调制机制显式建模系统元信息,有效应对多物理场混合挑战

大规模数据集已推动机器学习原子间势(MLIPs)在通用异相催化建模中的高精度发展。然而,由于训练数据存在空白,现有方法在处理自旋极化或高保真度系统时仍受限。为此,我们提出AQCat25,一个包含1350万条密度泛函理论(DFT)单点计算的互补数据集,旨在改善自旋极化与高保真度关键体系的建模能力。我们研究了将新数据集(如AQCat25)与更广泛的开放催化2020(OC20)数据集整合的方法,以构建具备自旋感知能力的模型,同时不牺牲通用性。结果表明,直接在AQCat25上微调通用模型会导致原始知识的灾难性遗忘;而联合训练策略则能有效提升新数据上的精度,且不损失整体性能。该联合方法带来新挑战:模型需处理混合保真度与混合物理(自旋极化/非极化)数据。我们证明,通过显式条件化模型对系统特定元数据的依赖(例如使用特征逐元素线性调制,FiLM),可成功应对这一挑战,并进一步提高模型精度。最终,本工作建立了一套有效的跨DFT保真度域融合协议,推动催化基础模型预测能力的提升。

原文摘要 · Abstract (English)

Large-scale datasets have enabled highly accurate machine learning interatomic potentials (MLIPs) for general-purpose heterogeneous catalysis modeling. There are, however, some limitations in what can be treated with these potentials because of gaps in the underlying training data. To extend these capabilities, we introduce AQCat25, a complementary dataset of 13.5 million density functional theory (DFT) single point calculations designed to improve the treatment of systems where spin polarization and/or higher fidelity are critical. We also investigate methodologies for integrating new datasets, such as AQCat25, with the broader Open Catalyst 2020 (OC20) dataset to create spin-aware models without sacrificing generalizability. We find that directly tuning a general model on AQCat25 leads to catastrophic forgetting of the original dataset's knowledge. Conversely, joint training strategies prove effective for improving accuracy on the new data without sacrificing general performance. This joint approach introduces a challenge, as the model must learn from a dataset containing both mixed-fidelity calculations and mixed-physics (spin-polarized vs. unpolarized). We show that explicitly conditioning the model on this system-specific metadata, for example by using Feature-wise Linear Modulation (FiLM), successfully addresses this challenge and further enhances model accuracy. Ultimately, our work establishes an effective protocol for bridging DFT fidelity domains to advance the predictive power of foundational models in catalysis.

机器学习势催化模拟自旋极化数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。