arXiv:2601.22076cs.LGcs.DC2026-01被引 13

剖析生成式AI推理能耗差异,揭示关键影响因素。

Where Do the Joules Go? Diagnosing Inference Energy Consumption

  • 通过46个模型、1858种配置的大规模实测,定位能耗根源。
  • 视频生成能耗可比图像高100倍,任务类型差超25倍。
  • 提出跨软硬件层的能耗诊断框架,适合能效优化研究者。

能源已成为机器学习计算的关键资源。尽管测量能耗并观察趋势是重要第一步,但准确理解与诊断差异成因对优化至关重要。为此,我们开展了一项大规模推理耗时与能耗测量研究,覆盖46个模型、7项任务及1,858种不同配置,使用NVIDIA H100和B200 GPU。实测发现:任务类型可导致能耗相差25倍,视频生成有时比图像生成高出100倍以上,GPU利用率差异亦能引起3至5倍能耗波动。基于这些发现,我们提出一个用于分析时间与能耗底层机制的框架。核心观点是:时间和能耗由内存、利用率等隐含指标决定,而这些指标又受算法、软件与硬件各层因素影响。该框架还可直接扩展至每瓦特吞吐量(throughput per watt),适用于电力受限的数据中心。

原文摘要 · Abstract (English)

Energy is now a critical ML computing resource. While measuring energy consumption and observing trends is a valuable first step, accurately understanding and diagnosing why those differences occur is crucial for optimization. To that end, we begin by presenting a large-scale measurement study of inference time and energy across the generative AI landscape with 46 models, 7 tasks, and 1,858 different configurations on NVIDIA H100 and B200 GPUs. Our empirical findings span order-of-magnitude variations: LLM task type can lead to 25$\times$ energy differences, video generation sometimes consumes more than 100$\times$ the energy of images, and GPU utilization differences can result in 3--5$\times$ energy differences. Based on our observations, we present a framework for reasoning about the underlying mechanisms that govern time and energy consumption. The essence is that time and energy are determined by latent metrics like memory and utilization, which are in turn affected by various factors across the algorithm, software, and hardware layers. Our framework also extends directly to throughput per watt, a critical metric for power-constrained datacenters.

能耗分析生成模型能效优化硬件感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。