用注意力嵌入优化超参数,自动平衡高性能计算的性能与功耗。
Attention-Informed Surrogates for Navigating Power-Performance Trade-offs in HPC
- 用作业遥测数据的注意力嵌入构建代理模型,捕捉性能动态。
- 在两个真实集群数据集上找到更优的运行时-功耗权衡解。
- 智能采样策略降低训练成本,提升结果稳定性,适合系统优化者。
高性能计算(HPC)调度器需在用户性能与设施资源限制间取得平衡,核心在于为任务选择最优节点数。本文提出一种基于代理模型的多目标贝叶斯优化(MOBO)框架,以自动化该决策过程。核心假设是:由作业遥测数据生成的注意力嵌入所指导的代理模型,能比传统回归方法更有效捕捉性能动态。同时,采用智能样本获取策略,确保方法的数据效率。在两个生产级HPC数据集上,该嵌入引导的方法始终生成质量更高的运行时-功耗帕累托前沿,且智能采样显著降低训练开销,提升结果稳定性。据我们所知,这是首个成功将嵌入引导代理模型应用于MOBO框架来联合优化真实工作负载下性能与功耗的研究。
原文摘要 · Abstract (English)
High-Performance Computing (HPC) schedulers must balance user performance with facility-wide resource constraints. The task boils down to selecting the optimal number of nodes for a given job. We present a surrogate-assisted multi-objective Bayesian optimization (MOBO) framework to automate this complex decision. Our core hypothesis is that surrogate models informed by attention-based embeddings of job telemetry can capture performance dynamics more effectively than standard regression techniques. We pair this with an intelligent sample acquisition strategy to ensure the approach is data-efficient. On two production HPC datasets, our embedding-informed method consistently identified higher-quality Pareto fronts of runtime-power trade-offs compared to baselines. Furthermore, our intelligent data sampling strategy drastically reduced training costs while improving the stability of the results. To our knowledge, this is the first work to successfully apply embedding-informed surrogates in a MOBO framework to the HPC scheduling problem, jointly optimizing for performance and power on production workloads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。