arXiv:2501.02600cs.DCcs.AI2025-01被引 80

针对云上大模型推理的温控与功耗问题,提出动态调度框架提升能效与稳定性。

TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms

  • 基于历史温电数据和SaaS弹性,动态调度推理任务
  • 降低热与功耗限流事件,提升系统效率
  • 适合云平台大模型服务部署与应急调度场景

生成式大语言模型(LLM)的兴起给云数据中心的温控与功耗管理带来挑战。传统方法难以应对LLM推理中毫秒级细粒度执行阶段带来的性能、温度与功耗差异。此外,模型并行、大小与量化等配置参数在性能、温度、功耗与输出质量间存在权衡。云环境常共置SaaS与IaaS工作负载,灵活性与可见性各异。本文提出TAPAS框架,专为云上LLM推理集群设计,增强冷却与功耗超分配能力,在降低总体拥有成本(TCO)的同时有效应对冷却与供电故障等紧急情况。系统利用历史温电数据及SaaS工作负载的可调性,实现:(1)在温电约束内高效部署新GPU虚拟机;(2)跨SaaS虚拟机路由推理请求;(3)动态重构虚拟机以应对负载突增与异常状况。在大规模GPU集群上的评估显示,显著减少热与功耗限流事件,提升系统效率。

原文摘要 · Abstract (English)

The rising demand for generative large language models (LLMs) poses challenges for thermal and power management in cloud datacenters. Traditional techniques often are inadequate for LLM inference due to the fine-grained, millisecond-scale execution phases, each with distinct performance, thermal, and power profiles. Additionally, LLM inference workloads are sensitive to various configuration parameters (e.g., model parallelism, size, and quantization) that involve trade-offs between performance, temperature, power, and output quality. Moreover, clouds often co-locate SaaS and IaaS workloads, each with different levels of visibility and flexibility. We propose TAPAS, a thermal- and power-aware framework designed for LLM inference clusters in the cloud. TAPAS enhances cooling and power oversubscription capabilities, reducing the total cost of ownership (TCO) while effectively handling emergencies (e.g., cooling and power failures). The system leverages historical temperature and power data, along with the adaptability of SaaS workloads, to: (1) efficiently place new GPU workload VMs within cooling and power constraints, (2) route LLM inference requests across SaaS VMs, and (3) reconfigure SaaS VMs to manage load spikes and emergency situations. Our evaluation on a large GPU cluster demonstrates significant reductions in thermal and power throttling events, boosting system efficiency.

大模型推理温控调度云平台优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。