提出评估能耗感知NAS基准的三大设计原则,揭示测量方法对结果的关键影响。
Guidelines for the Quality Assessment of Energy-Aware NAS Benchmarks
- 基于可靠测功、广泛GPU使用和全面能耗报告设计基准
- 发现Nvidia SMI API导致低功耗误报,与真实电表数据相关性差
- 提出校准方案,将工具误差从10.3%降至6.6%,适合能源效率研究者
神经架构搜索(NAS)虽加速了深度学习发展,但其搜索过程能耗日益增加。代理基准通过查询预训练代理模型估算性能,以降低全量训练成本。其中,能耗感知基准旨在实现模型精度与能耗之间的合理权衡。为此,本文提出三项设计原则:(i) 可靠的功耗测量,(ii) 广泛的GPU使用范围,(iii) 全面的成本报告。基于此分析EA-HAS-Bench发现,使用Nvidia SMI API会因采样率问题导致初始数据采集时出现异常低功耗估计,与外部电表实测值相关性差。实验显示单卡运行时功率范围为146W至305W,四卡并行时范围进一步缩小。为改善整体能耗报告,本文对Code Carbon等常用工具中的假设进行校准,使最大误差从10.3%降至8.9%(无先验负载),6.6%(有先验负载),从而为新型能耗感知基准的设计提供首套指导原则。
原文摘要 · Abstract (English)
Neural Architecture Search (NAS) accelerates progress in deep learning through systematic refinement of model architectures. The downside is increasingly large energy consumption during the search process. Surrogate-based benchmarking mitigates the cost of full training by querying a pre-trained surrogate to obtain an estimate for the quality of the model. Specifically, energy-aware benchmarking aims to make it possible for NAS to favourably trade off model energy consumption against accuracy. Towards this end, we propose three design principles for such energy-aware benchmarks: (i) reliable power measurements, (ii) a wide range of GPU usage, and (iii) holistic cost reporting. We analyse EA-HAS-Bench based on these principles and find that the choice of GPU measurement API has a large impact on the quality of results. Using the Nvidia System Management Interface (SMI) on top of its underlying library influences the sampling rate during the initial data collection, returning faulty low-power estimations. This results in poor correlation with accurate measurements obtained from an external power meter. With this study, we bring to attention several key considerations when performing energy-aware surrogate-based benchmarking and derive first guidelines that can help design novel benchmarks. We show a narrow usage range of the four GPUs attached to our device, ranging from 146 W to 305 W in a single-GPU setting, and narrowing down even further when using all four GPUs. To improve holistic energy reporting, we propose calibration experiments over assumptions made in popular tools, such as Code Carbon, thus achieving reductions in the maximum inaccuracy from 10.3 % to 8.9 % without and to 6.6 % with prior estimation of the expected load on the device.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。