arXiv:2504.06307cs.LGcs.AI2025-04被引 11

通过量化与本地推理降低大模型能耗,减碳超45%。

Optimizing Large Language Models: Metrics, Energy Efficiency, and Case Study Insights

  • 采用量化与本地推理技术优化大模型部署
  • 能耗与碳排放最高降低45%
  • 适合资源受限环境下的可持续AI应用

大型语言模型(LLMs)的快速普及带来了显著的能源消耗和碳排放,对生成式AI的可持续性构成严峻挑战。本文探讨了在LLM部署中集成节能优化技术以应对这些环境问题。我们提出了一个案例研究和框架,展示通过战略性量化和本地推理技术,可在不损害模型性能的前提下大幅降低LLM的碳足迹。实验结果表明,量化后能耗与碳排放最多可减少45%,特别适用于资源受限的环境。研究为实现人工智能可持续发展提供了可操作的洞察,同时保持高准确性和响应速度。

原文摘要 · Abstract (English)

The rapid adoption of large language models (LLMs) has led to significant energy consumption and carbon emissions, posing a critical challenge to the sustainability of generative AI technologies. This paper explores the integration of energy-efficient optimization techniques in the deployment of LLMs to address these environmental concerns. We present a case study and framework that demonstrate how strategic quantization and local inference techniques can substantially lower the carbon footprints of LLMs without compromising their operational effectiveness. Experimental results reveal that these methods can reduce energy consumption and carbon emissions by up to 45\% post quantization, making them particularly suitable for resource-constrained environments. The findings provide actionable insights for achieving sustainability in AI while maintaining high levels of accuracy and responsiveness.

大模型优化节能碳足迹量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。