arXiv:2502.05610cs.CL2025-02NAACL被引 28

首次系统评估大模型推理能耗,发现输出长度和批处理大小影响最大。

Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language Models

  • 测试多种模型、任务与提示对推理能耗的影响
  • 输出词数越多、响应越慢,耗能越高
  • 量化与优化批处理可显著节能,适合部署优化者

大语言模型在自然语言处理中表现出色,但其推理阶段的能耗问题长期被忽视。本研究首次系统性地在多种NLP任务中基准测试了大模型推理能耗,分析了模型、任务、提示及系统因素的影响。实验发现,推理能耗与生成输出长度和响应时间存在强相关性。通过量化、选择最优批量大小以及使用特定提示词,可显著降低能耗。该研究为模型部署中的能效优化提供了实证依据和实践建议。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly recognized for their exceptional generative capabilities and versatility across various tasks. However, the high inference costs associated with these models have not received adequate attention, particularly when compared to the focus on training costs in existing research. In response to this gap, our study conducts a comprehensive benchmarking of LLM inference energy across a wide range of NLP tasks, where we analyze the impact of different models, tasks, prompts, and system-related factors on inference energy. Specifically, our experiments reveal several interesting insights, including strong correlation of inference energy with output token length and response time. Also, we find that quantization and optimal batch sizes, along with targeted prompt phrases, can significantly reduce energy usage. This study is the first to thoroughly benchmark LLM inference across such a diverse range of aspects, providing insights and offering several recommendations for improving energy efficiency in model deployment.

大模型推理能耗能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。