arXiv:2505.06371cs.LGcs.AI2025-05NeurIPS被引 45

首个面向生成式AI推理能耗的开源基准测试工具

The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization

  • 构建真实服务环境下的推理能耗测量框架
  • 40个模型在6类任务中实测,优化可省超40%能耗
  • 适合关注绿色AI与系统能效的开发者

随着生成式AI在实际服务中的广泛应用,能源已成为关键瓶颈资源。然而,在构建机器学习系统时,能源仍常被忽视、探索不足或理解不深。本文提出ML.ENERGY基准测试套件和工具,用于在真实服务环境中测量推理能耗,并配套推出ML.ENERGY排行榜,为希望理解并优化生成式AI服务能耗的研究者提供宝贵资源。文中阐述了四个长期积累的机器学习能耗评估设计原则,并说明其在该基准中的实现方式。我们展示了2025年初版基准的成果,包括40种广泛使用的模型架构在6项不同任务中的能耗数据,揭示了机器学习设计选择对能耗的影响,并验证了自动化优化建议可在不改变模型计算内容的前提下,实现高达40%以上的能耗降低。该基准为开源项目,支持自定义模型与应用场景的扩展。

原文摘要 · Abstract (English)

As the adoption of Generative AI in real-world services grow explosively, energy has emerged as a critical bottleneck resource. However, energy remains a metric that is often overlooked, under-explored, or poorly understood in the context of building ML systems. We present the ML$.$ENERGY Benchmark, a benchmark suite and tool for measuring inference energy consumption under realistic service environments, and the corresponding ML$.$ENERGY Leaderboard, which have served as a valuable resource for those hoping to understand and optimize the energy consumption of their generative AI services. In this paper, we explain four key design principles for benchmarking ML energy we have acquired over time, and then describe how they are implemented in the ML$.$ENERGY Benchmark. We then highlight results from the early 2025 iteration of the benchmark, including energy measurements of 40 widely used model architectures across 6 different tasks, case studies of how ML design choices impact energy consumption, and how automated optimization recommendations can lead to significant (sometimes more than 40%) energy savings without changing what is being computed by the model. The ML$.$ENERGY Benchmark is open-source and can be easily extended to various customized models and application scenarios.

能耗测量生成式AI能效优化开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。