arXiv:2606.00401physics.comp-phcond-mat.mtrl-sci2026-06

用机器学习预测电子结构谱,加速上千原子系统计算。

Data-Driven Spectral Prediction for Accelerating Large-Scale Electronic Structure Calculations

  • 将谱预测转为切比雪夫多项式系数预测,突破大规模计算维度瓶颈。
  • 在2TB蛋白质二聚体数据上训练模型,显著减少自洽场迭代次数。
  • 适合需要加速量子化学模拟的高性能计算与材料设计研究者。

模拟包含数千个原子的大分子系统需要高度可扩展的方法。尽管现代密度泛函理论(DFT)代码已实现线性缩放,但在百亿亿次级架构上求解大型稀疏广义特征值问题仍是关键计算瓶颈。在LimitX项目背景下,我们提出一种数据驱动框架以加速此类计算。通过将机器学习目标从离散特征值改为插值切比雪夫多项式的系数,并对比全原子与片段化结构表示,成功克服了大规模谱预测的维度限制。我们在一个全新的2TB蛋白质二聚体数据集上训练了三种机器学习模型(核岭回归、图神经网络和随机森林)。预测得到的谱作为初始猜测,能有效跳过BigDFT中的早期自洽场(SCF)迭代。最终,这些谱预测器将部署用于动态优化即将推出的基于有理滤波的特征求解器(如正在初期开发的FrASE)。

原文摘要 · Abstract (English)

Simulating large molecular systems comprising thousands of atoms requires highly scalable methodologies. While modern Density Functional Theory (DFT) codes exhibit linear scaling, solving the associated large, sparse generalized eigenproblems remains a critical computational bottleneck on exascale architectures. In the context of the LimitX project, we propose a data-driven framework to accelerate these calculations. By shifting the machine learning target from discrete eigenvalues to the coefficients of an interpolating Chebyshev polynomial, and by comparing both all-atom and fragment-based structural representations, we successfully overcome the dimensionality constraints of large-scale spectral prediction. We investigate three machine learning models (Kernel Ridge Regression, Graph Neural Networks, and Random Forests) trained on a novel 2 TB dataset of protein dimers. The predicted spectra provide initial guesses that effectively bypass early Self-Consistent Field (SCF) iterations in BigDFT. Ultimately, these spectral predictors will be deployed to dynamically optimize upcoming rational filter-based eigensolvers, such as FrASE, which is currently in initial development.

电子结构机器学习量子化学加速计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。