arXiv:2410.11392physics.chem-phcs.LG2024-10被引 5

通过调整数据层级权重,用极少高精度数据实现精准激发能预测。

Investigating Data Hierarchies in Multifidelity Machine Learning for Excitation Energies

  • 引入随计算时间变化的动态缩放因子θ,优化多保真度数据利用效率。
  • 仅需2个高精度样本+大量低精度样本,即可达到高预测精度。
  • 提出Γ-曲线与误差轮廓图,直观展示模型误差来源与成本权衡。

近年来机器学习在量子化学计算中取得进展,使高精度计算更易获取。其中多保真度机器学习(MFML)方法通过结合不同精度的数据进行训练,通常采用固定缩放因子γ来平衡各保真度层级的样本数量,反映数据成本与稀疏性假设。本研究探讨了γ值变化对垂直激发能预测模型效率与精度的影响,并引入基于量子化学计算时间的新型缩放因子θ,使其随不同保真度下的计算耗时动态变化。提出一种新误差度量——MFML误差轮廓,全面展示各保真度层级对模型误差的贡献。结果表明,当使用较多低保真度样本时,仅需2个目标保真度样本即可实现高精度预测。进一步提出的Γ-曲线对比模型误差与数据生成时间成本,验证了多保真度模型能在极低训练成本下保持高准确性。

原文摘要 · Abstract (English)

Recent progress in machine learning (ML) has made high-accuracy quantum chemistry (QC) calculations more accessible. Of particular interest are multifidelity machine learning (MFML) methods where training data from differing accuracies or fidelities are used. These methods usually employ a fixed scaling factor, $γ$, to relate the number of training samples across different fidelities, which reflects the cost and assumed sparsity of the data. This study investigates the impact of modifying $γ$ on model efficiency and accuracy for the prediction of vertical excitation energies using the QeMFi benchmark dataset. Further, this work introduces QC compute time informed scaling factors, denoted as $θ$, that vary based on QC compute times at different fidelities. A novel error metric, error contours of MFML, is proposed to provide a comprehensive view of model error contributions from each fidelity. The results indicate that high model accuracy can be achieved with just 2 training samples at the target fidelity when a larger number of samples from lower fidelities are used. This is further illustrated through a novel concept, the $Γ$-curve, which compares model error against the time-cost of generating training samples, demonstrating that multifidelity models can achieve high accuracy while minimizing training data costs.

多保真度量子化学机器学习激发能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。