用vLLM构建真实场景下的大模型能效基准,助力绿色AI开发。
Benchmarking Energy Efficiency of Large Language Models Using vLLM
- 基于vLLM设计真实部署环境下的能效评测框架。
- 发现模型规模、架构和并发请求数显著影响推理能耗。
- 为开发者提供可落地的可持续AI系统优化参考。
大型语言模型(LLMs)的广泛应用正因部署与使用所需的大量能源而对气候产生日益增长的影响。为提高开发者在产品中部署LLMs时的能效意识,亟需获取更贴近实际生产场景的能效数据。现有研究虽已评估多种模型的能效,但往往无法反映真实部署条件。本文提出LLM Efficiency Benchmark,通过vLLM——一个高吞吐、生产就绪的LLM服务后端——模拟真实使用场景。我们分析了模型大小、架构及并发请求数对推理能效的影响。结果表明,该基准能更真实地反映实际部署条件,为致力于构建更可持续AI系统的开发者提供关键洞见。
原文摘要 · Abstract (English)
The prevalence of Large Language Models (LLMs) is having an growing impact on the climate due to the substantial energy required for their deployment and use. To create awareness for developers who are implementing LLMs in their products, there is a strong need to collect more information about the energy efficiency of LLMs. While existing research has evaluated the energy efficiency of various models, these benchmarks often fall short of representing realistic production scenarios. In this paper, we introduce the LLM Efficiency Benchmark, designed to simulate real-world usage conditions. Our benchmark utilizes vLLM, a high-throughput, production-ready LLM serving backend that optimizes model performance and efficiency. We examine how factors such as model size, architecture, and concurrent request volume affect inference energy efficiency. Our findings demonstrate that it is possible to create energy efficiency benchmarks that better reflect practical deployment conditions, providing valuable insights for developers aiming to build more sustainable AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。