实测大模型推理能耗,发现公开数据高估4-20倍,效率提升可降耗8-20倍。
Energy Use of AI Inference, Efficiency Pathways, and Test-Time Scaling
- 从吞吐量与节点功耗出发,构建真实部署下的能耗估算框架。
- 前沿模型每查询耗电中位数0.31瓦时,长查询场景升至3.91瓦时。
- 适合关注AI能效、数据中心规划及效率优化的研究者与工程师。
随着AI推理规模达数十亿次查询,单次查询的能耗估算对容量规划、效率干预和政策制定日益重要。然而,多数公开估算基于非生产环境,导致系统性高估。本文提出一种自下而上的框架,基于令牌吞吐量、节点功耗与开销,在大规模部署假设下估算推理能耗。对于参数量超过2000亿的前沿模型(部署于H100节点),我们估算中位能耗为0.31瓦时/查询(四分位距0.16-0.60),表明广泛引用的估算值被高估了4-20倍。在测试时扩展场景(比典型查询长15倍)下,中位能耗上升13倍至3.91瓦时(四分位距2.15-7.05)。跨模型、服务系统与硬件,我们估计存在8-20倍的直接节能空间。在数据中心层面,日均处理10亿次查询需0.7吉瓦时;若10%为长查询,能耗升至1.7吉瓦时/日;通过效率干预,能耗可降至0.8吉瓦时/日,有效缓解测试时扩展带来的能源压力。
原文摘要 · Abstract (English)
As AI inference scales to billions of queries, estimates of per-query energy use are increasingly important for capacity planning, efficiency interventions, and policy. Yet many public estimates assume non-production settings, leading to systematic overestimation. We introduce a bottom-up framework estimating inference energy from token throughput, node power, and overhead under large-scale deployment assumptions. For frontier-scale models (>200B parameters) on H100 nodes, we estimate a median energy of 0.31 Wh/query (IQR 0.16-0.60), indicating widely cited estimates are overstated by 4-20x. In test-time scaling scenarios 15x longer than typical queries, the median energy rises 13x to 3.91 Wh (IQR 2.15-7.05). Across models, serving systems, and hardware, we estimate 8-20x line-of-sight energy reductions. At datacenter scale, serving 1 billion queries/day requires 0.7 GWh; if 10% are long queries, demand rises to 1.7 GWh/day. With efficiency interventions, it falls to 0.8 GWh/day, mitigating the energy impact of test-time scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。