arXiv:2511.05597cs.AIcs.LG2025-11被引 24

测量大模型推理能耗,揭示影响能效的关键因素。

From Prompts to Power: Measuring the Energy Footprint of LLM Inference

  • 通过3.25万次测量,分析21种GPU与155种模型的能效差异。
  • 发现模型架构和运行配置对能耗影响显著,误差小于10%。
  • 推出浏览器插件,实时提示生成式AI的碳足迹。

大型语言模型(LLMs)的快速扩展带来了前所未有的能源需求,其消耗已从训练阶段延伸至大规模推理任务,后者常占全生命周期能耗的主导地位。部署这些模型需依赖高功耗的GPU基础设施,部分数据中心甚至计划采用核能供电。然而,针对推理阶段能源消耗的系统性研究仍显不足。本文基于超过32,500次测量,涵盖21种GPU配置与155种模型架构(从小型开源模型到前沿系统),利用vLLM推理引擎,在提示级别量化了能源使用情况,并识别出架构与运行因素如何影响能耗。基于这些发现,我们构建了一个可准确预测未见模型与硬件组合下推理能耗的预测模型,并将其集成为浏览器扩展,以提升公众对生成式AI环境影响的认知。

原文摘要 · Abstract (English)

The rapid expansion of Large Language Models (LLMs) has introduced unprecedented energy demands, extending beyond training to large-scale inference workloads that often dominate total lifecycle consumption. Deploying these models requires energy-intensive GPU infrastructure, and in some cases has even prompted plans to power data centers with nuclear energy. Despite this growing relevance, systematic analyses of inference energy consumption remain limited. In this work, we present a large-scale measurement-based study comprising over 32,500 measurements across 21 GPU configurations and 155 model architectures, from small open-source models to frontier systems. Using the vLLM inference engine, we quantify energy usage at the prompt level and identify how architectural and operational factors shape energy demand. Building on these insights, we develop a predictive model that accurately estimates inference energy consumption across unseen architectures and hardware, and implement it as a browser extension to raise awareness of the environmental impact of generative AI.

大模型能耗测量绿色AI推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。