arXiv:2606.16226cs.LG2026-06

用主动学习与生成学习预测并行化学计算的运行时参数

Prediction of Runtime Parameters of Parallel Chemistry Applications via Active and Generative Learning

论文配图:Prediction of Runtime Parameters of Parallel Chemistry Applications via Active and Generative Learning
图 1 · 摘自论文原文
  • 结合主动学习与生成学习,优化梯度提升回归模型
  • 在量子化学计算中实现0.023%的平均绝对误差,决定系数达99.9%
  • 仅需原始数据20-25%即可达到0.2%的误差,适合数据稀缺场景

本文提出两种基于机器学习的方案,用于预测高度可扩展的并行化学计算的运行时参数。方法结合主动学习与生成学习,并从多种机器学习模型中选择经验确定的梯度提升回归树模型。在耦合簇单双激发(Coupled-Cluster with Singles and Doubles)计算上的评估显示,模型的平均绝对误差百分比(MAPE)低至0.023%,决定系数高达99.9%。此外,结合主动学习缓解训练数据不足问题后,仅使用原始数据的20%-25%时,模型仍能达到约0.2%的MAPE。

原文摘要 · Abstract (English)

In this work, we develop two main Machine Learning based approaches to predict the runtime parameters of highly scalable parallel chemistry computations.These approaches employ active and generative learning together with the empirically determined gradient boosted regression tree models chosen among a rich suite of machine learning models. When evaluated on Coupled-Cluster with Singles and Doubles computations, our models achieve a mean absolute error percentage (MAPE) as low as 0.023 and a coefficient of determination as high as 99.9%. Furthermore, when combined with active learning to mitigate the lack of large amounts of training data, our models score a MAPE about 0.2 with 20-25% of the original dataset.

机器学习化学计算主动学习预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。