利用分子哈密顿矩阵数据提升机器学习势能模型性能
Learning from the electronic structure of molecules across the periodic table
- 用哈密顿矩阵中轨道相互作用数据训练新模型
- 在58种元素、150个原子规模下实现高精度预测
- 适合小样本场景的通用势能模型研究者
机器学习原子间势能(MLIPs)需大量原子结构数据来学习力和能量,其性能随训练集增大持续提升。然而,这些数据背后的哈密顿矩阵H中蕴含的更大量信息至今未被利用。本文提出一种将轨道相互作用数据融入训练流程的方法。首先推出新型哈密顿矩阵预测模型HELM,可处理含100+原子、高元素多样性及包含弥散函数的大基组结构。为此构建了全新数据集OMol_CSH_58k,涵盖58种元素、最大150个原子、基组为def2-TZVPD。进一步提出“哈密顿预训练”方法,从有限原子结构中提取有意义的原子环境描述符,并复用该共享嵌入空间,显著提升低数据条件下的能量预测性能。结果表明,电子相互作用是表征化学空间的丰富且可迁移的数据源。
原文摘要 · Abstract (English)
Machine-Learned Interatomic Potentials (MLIPs) require vast amounts of atomic structure data to learn forces and energies, and their performance continues to improve with training set size. Meanwhile, the even greater quantities of accompanying data in the Hamiltonian matrix H behind these datasets has so far gone unused for this purpose. Here, we provide a recipe for integrating the orbital interaction data within H towards training pipelines for atomic-level properties. We first introduce HELM ("Hamiltonian-trained Electronic-structure Learning for Molecules"), a state-of-the-art Hamiltonian prediction model which bridges the gap between Hamiltonian prediction and universal MLIPs by scaling to H of structures with 100+ atoms, high elemental diversity, and large basis sets including diffuse functions. To accompany HELM, we release a curated Hamiltonian matrix dataset, 'OMol_CSH_58k', with unprecedented elemental diversity (58 elements), molecular size (up to 150 atoms), and basis set (def2-TZVPD). Finally, we introduce 'Hamiltonian pretraining' as a method to extract meaningful descriptors of atomic environments even from a limited number atomic structures, and repurpose this shared embedding space to improve performance on energy-prediction in low-data regimes. Our results highlight the use of electronic interactions as a rich and transferable data source for representing chemical space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。