arXiv:2604.09048cs.DCcs.AI2026-04被引 4

构建首个开源大模型能效基准,助力绿色部署。

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures

论文配图:Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
图 1 · 摘自论文原文
  • 建立跨10款GPU的5000+次实验能效数据集
  • 实测显示服务器场景节能最高达70%
  • 适合关注大模型低碳部署的研究与工程人员

尽管大语言模型(LLMs)的高能耗已广受关注,但系统运维者缺乏针对异构硬件能量权衡的能效部署指导,原因在于缺少能效感知的基准和数据。本文提出Watt Counts:目前最大的开源大模型能效数据集,包含50个LLMs在10款NVIDIA GPU上进行批量与服务器场景下的5000+次实验,配套可复现的开源基准,支持社区持续扩展。基于该数据集,我们开展系统级研究发现,GPU选型对能效影响显著,最优硬件因模型和部署场景而异,凸显异构系统中硬件感知部署的关键性。基于此,我们证明在服务器场景下可实现最高70%的能耗降低,且对用户体验影响极小;批量场景下也能降低20%能耗。

原文摘要 · Abstract (English)

While the large energy consumption of Large Language Models (LLMs) is recognized by the community, system operators lack guidance for energy-efficient LLM inference deployments that leverage energy trade-offs of heterogeneous hardware due to a lack of energy-aware benchmarks and data. In this work we address this gap with Watt Counts: the largest open-access dataset of energy consumption of LLMs, with over 5,000 experiments for 50 LLMs across 10 NVIDIA Graphics Processing Units (GPUs) in batch and server scenarios along with a reproducible, open-source benchmark that enables community submissions to expand this dataset. Leveraging this dataset, we conduct a system-level study of LLM inference across heterogeneous GPU architectures and show that GPU selection is crucial for energy efficiency outcomes and that optimal hardware choices vary significantly across models and deployment scenarios, demonstrating the critical importance of hardware-aware deployment in heterogeneous LLM systems. Guided by our data and insights, we show that practitioners can reduce energy consumption by up to 70% in server scenarios with negligible impact on user experience, and by up to 20% in batch scenarios.

能效优化大模型部署异构计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。