提出新调度框架,同时降低大模型推理的碳排放、用水和能耗。
Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
- 用机器学习优化跨地域数据中心的LLM调度策略。
- 每20-50次请求耗水500毫升,推理阶段碳足迹是训练的25倍。
- 适合关注绿色AI与可持续云服务的研究者与工程师。
近年来,大型语言模型(如ChatGPT、CoPilot、Gemini)在各领域广泛应用。尽管研究多聚焦于降低模型训练开销,但其推理阶段的环境影响日益突出。最新研究显示,推理阶段的运营成本每年可超过训练成本的25倍,累积碳足迹也远超训练阶段。此外,每20至50次推理请求需消耗约500毫升淡水。为应对这些可持续性挑战,本文提出新型框架SLIT,协同优化大模型服务质量(首次令牌响应时间)、碳排放、用水量及能源成本。该框架采用基于机器学习的元启发式算法,在地理分布的云数据中心中提升大模型部署的可持续性。随着大模型普及,此类技术愈发关键。
原文摘要 · Abstract (English)
In recent years, Large Language Models (LLM) such as ChatGPT, CoPilot, and Gemini have been widely adopted in different areas. As the use of LLMs continues to grow, many efforts have focused on reducing the massive training overheads of these models. But it is the environmental impact of handling user requests to LLMs that is increasingly becoming a concern. Recent studies estimate that the costs of operating LLMs in their inference phase can exceed training costs by 25x per year. As LLMs are queried incessantly, the cumulative carbon footprint for the operational phase has been shown to far exceed the footprint during the training phase. Further, estimates indicate that 500 ml of fresh water is expended for every 20-50 requests to LLMs during inference. To address these important sustainability issues with LLMs, we propose a novel framework called SLIT to co-optimize LLM quality of service (time-to-first token), carbon emissions, water usage, and energy costs. The framework utilizes a machine learning (ML) based metaheuristic to enhance the sustainability of LLM hosting across geo-distributed cloud datacenters. Such a framework will become increasingly vital as LLMs proliferate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。