从碳排放角度优化大模型服务,兼顾运行与制造阶段的环境影响。
Towards Sustainable Large Language Model Serving
- 分析不同参数量模型在两类GPU上的能效与碳排放表现。
- 量化了运行碳排放(基于电网碳强度)与制造碳排放(基于芯片面积和内存)。
- 揭示同时考虑运行与制造碳排可显著提升大模型服务可持续性。
本文从碳排放视角研究大语言模型,涵盖运行与制造两种碳排放类型,为可持续大模型服务提供路径。我们使用两代Nvidia GPU(RTX6000 Ada与T4)对1B、3B、7B参数量的LLaMA模型进行性能与能耗表征。基于三个电网区域的碳强度数据,建立运行碳排放模型;结合芯片面积与内存大小,构建制造碳排放模型。该分析使我们深入理解大模型服务中的性能、能耗与碳排放关系。结果表明,通过同时优化运行与制造碳排放,可有效提升大模型服务的可持续性。
原文摘要 · Abstract (English)
In this work, we study LLMs from a carbon emission perspective, addressing both operational and embodied emissions, and paving the way for sustainable LLM serving. We characterize the performance and energy of LLaMA with 1B, 3B, and 7B parameters using two Nvidia GPU types, a latest-generation RTX6000 Ada and an older-generation T4. We analytically model operational carbon emissions based on energy consumption and carbon intensities from three grid regions -- each representing a different energy source mix, and embodied carbon emissions based on chip area and memory size. Our characterization and modeling provide us with an in-depth understanding of the performance, energy, and carbon emissions of LLM serving. Our findings highlight the potential for optimizing sustainable LLM serving systems by considering both operational and embodied carbon emissions simultaneously.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。