用一个模型实现主族元素分子性质的高精度预测,且可外推到更大体系。
Coupled-cluster molecular properties across the main group that extrapolate beyond training size

- 基于单个DFT计算预测有效哈密顿量,再推导多种电子结构性质。
- 在九种主族元素上误差比传统泛函降低3.8至230倍,大分子误差仅2%。
- 通过哈密顿量建模实现正确缩放性,可外推至58原子链,适合大分子研究。
耦合簇理论是分子电子结构性质的精度标准,但计算成本过高,而密度泛函理论虽廉价却存在系统偏差。本文提出MEHnet-MG模型,仅需一次低成本的B3LYP/def2-SVP计算即可预测有效一电子哈密顿量,并从中推导能量、光学带隙、偶极矩、四极矩、极化率、Mulliken原子电荷和Mayer键级等多种性质,达到耦合簇精度,覆盖九种主族元素(含磷、硫、氯等欠覆盖元素)。模型基于自建多性质数据集训练,该数据集在CCSD(T)/cc-pVTZ水平计算。在独立测试集上,其所有性质误差相较半局域、杂化及双杂化DFT降低3.8至230倍(以复合CCSD(T)/cc-pVTZ为参考),每分子仅增加约25毫秒推理时间,实现耦合簇级精度的快速预测。关键在于从预测哈密顿量出发而非堆叠原子特征,使模型天然具备正确的尺寸标度:在π共轭寡噻吩类分子中,对44原子和37原子体系的有限场CCSD极化率与EOM-CCSD光学带隙匹配精度达~2%,并成功外推至58原子链,而基于特征池化的架构因结构限制无法实现此类外推。因此,准确外推由模型归纳偏置决定,而非训练数据规模。
原文摘要 · Abstract (English)
Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine application, whereas density-functional theory is affordable yet systematically biased. We resolve this trade-off with a single equivariant network, MEHnet-MG, that predicts an effective one-electron Hamiltonian from one inexpensive B3LYP/def2-SVP calculation and derives a broad suite of properties from it (energy, optical gap, dipole, quadrupole, polarizability, Mulliken atomic charges, and Mayer bond orders) at coupled-cluster accuracy across nine main-group elements, including the under-served phosphorus, sulfur, and chlorine chemistries. The model is trained on a new in-house dataset of multi-property labels computed at the CCSD(T) level for all nine elements. On a held-out test set, it reduces the error of every property by a factor of 3.8 to 230 relative to semi-local, hybrid, and double-hybrid DFT (referenced to composite CCSD(T)/cc-pVTZ; Methods), while adding only ~25 ms wall time per molecule, delivering coupled-cluster-quality predictions at the cost of a single DFT calculation. Critically, deriving every property from a predicted Hamiltonian rather than pooling per-atom features builds the correct size-scaling into the model architecture: on pi-conjugated oligothiophenes it matches finite-field CCSD polarizability and the EOM-CCSD optical gap to ~2% at the largest sizes where those references remain affordable (44 and 37 atoms, where a single CCSD field point already costs ~500x the model's entire inference) and extrapolates the corrected trends to 58-atom chains, a regime where pooling-based architectures fail by construction. Accurate extrapolation is therefore set by the model's inductive bias rather than by the training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。