arXiv:2606.00794cond-mat.mtrl-scics.LG2026-06

构建5万条数据集,用机器学习加速二维材料催化研究

Benchmark Dataset for Catalysis on 2D MXenes

论文配图:Benchmark Dataset for Catalysis on 2D MXenes
图 1 · 摘自论文原文
  • 用5万条第一性原理数据训练机器学习势函数
  • 预测精度达±10meV/A力与±1meV/原子能量
  • 适合材料催化设计与大规模模拟研究者

结合第一性原理计算与机器学习,旨在加速新型材料催化行为的探索。研究聚焦于二维Ti₂CTy MXene,其丰富的表面化学使其成为催化理想候选。由于真实条件下结构与组成计算成本过高,常规密度泛函理论(DFT)难以胜任。为此,我们生成了包含5万条DFT计算用于训练、1万条用于测试的综合性数据集,涵盖Ti₂CTy MXene构型与分子体系,并额外构建1000个全新大型系统测试集,评估模型泛化能力。训练并验证了包括EquiformerV2、MACE、MatRIS和UPET在内的主流机器学习原子间势模型,可准确预测原子受力与形成能——这些是进行结构与催化研究时需反复计算的关键量。该结合框架在CPU上实现约10³~4×10³倍加速,同时保持高精度(力误差约±10 meV/A,每原子能量误差约±1 meV),为更高效地研究MXene催化行为铺平道路。此外,我们对模型进行了广泛定性评估,强调基于仿真结果的全面比较超越传统基准指标的重要性。数据集与训练模型及代码已公开于https://huggingface.co/datasets/CatalystAnonymous/catalyst_mxenes。

原文摘要 · Abstract (English)

Merging first-principles calculations with machine learning (ML), we aim to accelerate the exploration of catalytic behaviour in novel materials. We focus on two-dimensional (2D) Ti$_2$CT$_y$ MXenes, whose versatile surface chemistry makes them particularly compelling candidates for catalysis. Resolving their composition and structure under realistic conditions exceeds the reach of standard density functional theory (DFT) due to computational cost. To address this challenge, we generate a comprehensive dataset of 50,000 DFT calculations for training and 10,000 for testing, encompassing both Ti$_2$CT$_y$ MXene configurations and molecular systems, along with an additional test dataset with 1000 genuinely new, larger systems to investigate how well models generalise. We train and validate widely used and competitive machine learning interatomic potential (MLIP) models, including EquiformerV2, MACE, MatRIS, and UPET, that accurately predict atomic forces and formation energies -- quantities that DFT must repeatedly compute for structural and catalytic investigations -- for these 2D materials. This combined DFT-ML framework achieves computational acceleration on the order of approximately $1-4 \cdot 10^3$ (on a CPU) while maintaining desired-level accuracy (approximately +/- $10$ meV/A for forces and approximately +/- $1$ meV for per-atom energies), paving the way for more efficient investigations of MXene catalytic behaviour. Moreover, we perform an extensive qualitative evaluation of the trained models, showcasing the importance of comprehensive simulation-based comparison beyond benchmark metrics. The dataset and the trained models with the code are available at https://huggingface.co/datasets/CatalystAnonymous/catalyst_mxenes.

催化材料机器学习势二维材料第一性原理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。