arXiv:2505.09598cs.CYcs.AI2025-05被引 147

量化大模型推理的能耗水耗碳排,揭示其环境代价

How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference

  • 结合公开接口数据与厂商环境因子,推断硬件配置并评估能耗
  • 最耗能模型每长提示耗电超29瓦时,是高效系统的65倍以上
  • 适合关注AI可持续性、绿色计算及政策制定的研究者与企业

本文提出一种面向基础设施的大规模语言模型(LLM)推理环境足迹评估框架,涵盖30个先进商用模型。该框架融合公开API性能数据、企业特定环境乘数及硬件配置的统计推断,并采用交叉效率数据包络分析(DEA)对模型按性能-环境成本比进行排名。同时提供动态更新仪表盘,可视化各模型的能源、水和碳排放指标。结果显示,最耗能模型每处理一次长提示耗电超过29瓦时,是最高效率系统的65倍以上;即便单次短查询仅耗电0.42瓦时,若日均调用7亿次,年耗电量相当于3.5万美国家庭,蒸发淡水达120万人年饮用需求,碳排放需一座芝加哥大小的森林才能抵消。这些发现凸显一个悖论:随着AI变得更便宜更快,全球普及反而导致资源消耗不成比例增长。本方法为人工智能部署的可持续性评估与问责提供了标准化、实证基础。

原文摘要 · Abstract (English)

This paper introduces an infrastructure-aware benchmarking framework for quantifying the environmental footprint of LLM inference across 30 state-of-the-art models in commercial datacenters. The framework combines public API performance data with company-specific environmental multipliers and statistical inference of hardware configurations. We additionally utilize cross-efficiency Data Envelopment Analysis (DEA) to rank models by performance relative to environmental cost and provide a dynamically updated dashboard that visualizes model-level energy, water, and carbon metrics. Results show the most energy-intensive models exceed 29 Wh per long prompt, over 65 times the most efficient systems. Even a 0.42 Wh short query, when scaled to 700M queries/day, aggregates to annual electricity comparable to 35{,}000 U.S. homes, evaporative freshwater equal to the annual drinking needs of 1.2M people, and carbon emissions requiring a Chicago-sized forest to offset. These findings highlight a growing paradox: as AI becomes cheaper and faster, global adoption drives disproportionate resource consumption. Our methodology offers a standardized, empirically grounded basis for sustainability benchmarking and accountability in AI deployment.

大模型环境足迹可持续性能效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。