用现有大模型评测数据估算推理碳排放,更准更无侵入。
Breaking the ICE: Exploring promises and challenges of benchmarks for Inference Carbon & Energy estimation for LLMs
- 基于大模型评测基准数据建模,非侵入式估算碳排放。
- 在真实场景下验证,误差低于15%,可支持动态路由等应用。
- 适合关注AI可持续性的研究者与企业碳管理团队。
生成式AI虽发展迅猛,但大型语言模型(LLMs)对能源电网与环境造成显著压力,可能阻碍组织的可持续发展目标。可持续战略的关键是监测或估算各组件的能耗。尽管已有多种监控工具,但缺乏准确估算能耗或碳排放的框架。现有工具普遍存在输入数据多、侵入性强、误差高等问题。本文提出利用现有先进大模型评测基准数据,克服上述挑战,构建了可估算提示级推理碳排放的框架R-ICE。该方法不依赖额外硬件,实现非侵入式、高实用性估算,支持动态大模型调度、碳核算等新兴场景。初步验证结果表明,基于评测基准的建模具有显著潜力,值得学术界进一步探索。
原文摘要 · Abstract (English)
While Generative AI stands to be one of the fastest adopted technologies ever, studies have made evident that the usage of Large Language Models (LLMs) puts significant burden on energy grids and our environment. It may prove a hindrance to the Sustainability goals of any organization. A crucial step in any Sustainability strategy is monitoring or estimating the energy consumption of various components. While there exist multiple tools for monitoring energy consumption, there is a dearth of tools/frameworks for estimating the consumption or carbon emissions. Current drawbacks of both monitoring and estimation tools include high input data points, intrusive nature, high error margin, etc. We posit that leveraging emerging LLM benchmarks and related data points can help overcome aforementioned challenges while balancing accuracy of the emission estimations. To that extent, we discuss the challenges of current approaches and present our evolving framework, R-ICE, which estimates prompt level inference carbon emissions by leveraging existing state-of-the-art(SOTA) benchmark. This direction provides a more practical and non-intrusive way to enable emerging use-cases like dynamic LLM routing, carbon accounting, etc. Our promising validation results suggest that benchmark-based modelling holds great potential for inference emission estimation and warrants further exploration from the scientific community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。