arXiv:2605.13496cs.DCcs.LG2026-05

用博弈强化学习优化大模型推理,显著降低云数据中心的能耗与碳排。

MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters

论文配图:MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters
图 1 · 摘自论文原文
  • 设计多智能体博弈强化学习框架,协同优化响应速度与环境成本。
  • 相比现有方案,降低18%响应延迟、33%碳排放、43%用水量、11%能耗。
  • 适合关注绿色计算与高效推理部署的研究者与工程师。

大语言模型(LLM)在云平台中的应用日益广泛,其推理阶段占整个生命周期能源消耗的高达90%,远超训练阶段。随着推理请求量上升,对环境的影响愈发显著,尤其体现在碳排放和水资源消耗方面。为提升云数据中心中LLM推理的可持续性,本文提出一种新型多智能体博弈强化学习框架MARLIN,旨在协同优化时延(TTFT)、碳排放、用水量及能源成本。实验表明,相较于当前最先进的推理管理框架,MARLIN在各项指标上均实现显著改善:TTFT降低至少18%,碳排放减少33%,用水量下降43%,能源成本降低11%。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become increasingly prevalent in cloud-based platforms, propelled by the introduction of AI-based consumer and enterprise services. LLM inference requests in particular account for up to 90% of total LLM lifecycle energy use, dwarfing training energy costs. The rising volume of LLM inference requests is increasing environmental footprints, particularly carbon emissions and water consumption. To improve sustainability for LLM inference serving in cloud datacenter environments, we propose a novel multi-agent game-theoretic reinforcement learning framework called MARLIN to co-optimize time-to-first token (TTFT), carbon emissions, water usage, and energy costs associated with LLM inference. MARLIN demonstrates a reduction of at least 18% in TTFT, 33% in carbon emissions, 43% in water usage, and 11% in energy costs compared to state-of-the-art LLM inference management frameworks.

大模型推理绿色计算强化学习可持续性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。