arXiv:2509.20667cs.LGcs.CE2025-09被引 1

用机器学习预测化学计算资源消耗,帮用户省钱省时间。

Guiding Application Users via Estimation of Computational Resources for Massively Parallel Chemistry Computations

  • 用梯度提升模型预测超算上化学计算的运行时间
  • 在前沿和极光超算上误差分别低至2.3%和7.3%
  • 小样本下用主动学习仅450次实验就达20%误差

本文开发基于机器学习的策略,预测大规模并行化学计算(如耦合簇方法)所需的资源成本,帮助用户在提交昂贵超算任务前做出决策。通过预测应用执行时间,可确定最优运行参数,如节点数和分块大小。针对用户两个核心问题:一是最短时间配置,即给定问题规模和目标超算,找到使执行时间最短的参数组合;二是最低开销配置,即在给定问题规模下,最小化节点时数的节点数与分块大小。我们评估了多种机器学习模型,基于在能源部前沿(Frontier)和极光(Aurora)超算上运行CCSD程序的大量运行参数数据。实验表明,预测单次CCSD迭代总执行时间时,梯度提升(GB)模型在极光和前沿上的平均绝对百分比误差(MAPE)分别为0.073和0.023。当难以通过大量实验收集数据时,主动学习仅需约450次实验即可达到约0.2的MAPE。

原文摘要 · Abstract (English)

In this work, we develop machine learning (ML) based strategies to predict resources (costs) required for massively parallel chemistry computations, such as coupled-cluster methods, to guide application users before they commit to running expensive experiments on a supercomputer. By predicting application execution time, we determine the optimal runtime parameter values such as number of nodes and tile sizes. Two key questions of interest to users are addressed. The first is the shortest-time question, where the user is interested in knowing the parameter configurations (number of nodes and tile sizes) to achieve the shortest execution time for a given problem size and a target supercomputer. The second is the cheapest-run question in which the user is interested in minimizing resource usage, i.e., finding the number of nodes and tile size that minimizes the number of node-hours for a given problem size. We evaluate a rich family of ML models and strategies, developed based on the collections of runtime parameter values for the CCSD (Coupled Cluster with Singles and Doubles) application executed on the Department of Energy (DOE) Frontier and Aurora supercomputers. Our experiments show that when predicting the total execution time of a CCSD iteration, a Gradient Boosting (GB) ML model achieves a Mean Absolute Percentage Error (MAPE) of 0.023 and 0.073 for Aurora and Frontier, respectively. In the case where it is expensive to run experiments just to collect data points, we show that active learning can achieve a MAPE of about 0.2 with just around 450 experiments collected from Aurora and Frontier.

机器学习超算优化化学计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。