研究大模型生成策略如何影响显卡能耗,发现选对方法能省电且不降质量。
Energy-Conscious LLM Decoding: Impact of Text Generation Strategies on GPU Energy Consumption
- 对比多种生成策略在不同任务下的能耗表现。
- 相同质量下,不同策略能耗差可达数倍。
- 为节能应用提供无需牺牲质量的选型参考。
解码策略显著影响大语言模型(LLMs)生成文本的质量与多样性,但其对计算资源(尤其是GPU能耗)的影响尚未得到充分研究。本文探讨了文本生成解码技术与能效之间的关系,重点关注不同任务和配置下生成质量与GPU能耗之间的权衡。通过在翻译、数学问题求解、编程和开放式文本生成等多类任务上基准测试多种策略,揭示了选择合适的解码方法及其调优超参数,不仅影响输出质量,还对能耗有显著影响。研究发现,解码策略的选择可大幅改变GPU能耗,即使对输出质量影响甚微。不同策略在质量与能效之间存在权衡,不存在适用于所有指标的最优方案。据我们所知,这是首个从能耗角度系统分析LLM解码策略的研究,为构建高质量且节能的应用提供了实用洞见。
原文摘要 · Abstract (English)
Decoding strategies significantly influence the quality and diversity of the generated text in Large Language Models (LLMs), yet their impact on computational resources, particularly GPU energy consumption, is insufficiently studied. This paper investigates the relationship between text generation decoding techniques and energy efficiency, focusing on the trade-off between generation quality and GPU energy usage across diverse tasks and decoding configurations. By benchmarking multiple strategies across various tasks, including Translation, Math Problem Solving, Coding, and Open-ended text generation, we reveal how selecting appropriate decoding techniques with their tuned hyperparameters affects text quality and has measurable implications for energy consumption. Our findings show that the choice of decoding strategy can greatly impact GPU energy usage, even when it has a minimal effect on output quality. Different strategies also involve trade-offs between quality and energy efficiency, and no single decoding method is best in all cases across every metric. To the best of our knowledge, this is one of the first studies to examine decoding strategies in LLMs from the perspective of energy consumption, providing useful insights for building energy-efficient applications without compromising text generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。