arXiv:2605.02965cs.LGcs.SY2026-05

用扩散模型增强奖励信号,优化AI内容生成的能耗与调度效率。

Joint Energy Management and Coordinated AIGC Workload Scheduling for Distributed Data Centers: A Diffusion-Aided Reward Shaping Approach

论文配图:Joint Energy Management and Coordinated AIGC Workload Scheduling for Distributed Data Centers: A Diffusion-Aided Reward Shaping Approach
图 1 · 摘自论文原文
  • 引入扩散模型生成补充奖励,缓解强化学习中的奖励稀疏问题。
  • 联合优化多厂商AIGC任务调度与数据中心能耗,提升系统整体收益。
  • 适用于需兼顾成本、质量与异构模型的AI内容生成服务场景。

人工智能生成内容(AIGC)已成为自动化生成多样化定制内容的变革性范式,导致云数据中心计算负载急剧增长。AIGC服务提供商(ASPs)亟需战略性调度任务以降低能耗并保障内容质量。然而,AIGC服务具有模型异构、服务质量隐式评估及推理过程复杂等特性,带来严峻挑战。为此,我们提出一种联合能源管理与协同AIGC工作负载调度框架,通过显式数学表征服务质量,促进跨ASP任务迁移与细粒度推理配置,并综合考虑数据中心内多种能源资源以增强用电灵活性。进一步构建系统效用最大化问题,平衡服务收益与运营惩罚及成本。然而,任务调度决策间的强耦合导致奖励稀疏,制约现有深度强化学习(DRL)算法效果。为此,我们设计了一种扩散模型辅助的奖励塑形方法,通过多步去噪过程合成互补奖励信号,无缝集成于DRL中,实现稀疏反馈下调度策略的高效学习。基于真实模型与数据集的实验表明,该方案能有效应对电价波动与AIGC模型异构性,相比基准方法显著提升学习收敛速度与系统效用。

原文摘要 · Abstract (English)

Artificial intelligence-generated content (AIGC) has emerged as a transformative paradigm for automating the creation of diverse and customized content, giving rise to rapidly growing computational workloads in cloud data centers. It is imperative for AIGC service providers (ASPs) to strategically schedule AIGC workloads to reduce data center energy costs while guaranteeing high-quality content generation. However, the distinctive characteristics of AIGC services pose critical challenges, including model heterogeneity across ASPs, implicit service quality evaluation, and complex inference process control. To tackle these challenges, we propose a joint energy management and coordinated AIGC workload scheduling framework, which introduces an explicit mathematical characterization of service quality to promote both job transfer among ASPs and fine-grained inference process configuration. Moreover, various energy resources within data centers are jointly considered to enhance power usage flexibility. Subsequently, a system utility maximization problem is formulated to balance AIGC service revenue with operational penalties and costs. Nevertheless, the strong coupling among job scheduling decisions induces severe reward sparsity, which limits the effectiveness of existing deep reinforcement learning (DRL) algorithms. To address this issue, we develop a diffusion model-aided reward shaping approach to synthesize complementary reward signals through a multi-step denoising process. This approach is seamlessly integrated with DRL to enable efficient learning of scheduling policies under sparse environmental feedback. Experiments based on real-world models and datasets demonstrate that our scheme effectively accommodates electricity price fluctuations and AIGC model heterogeneity, while achieving superior learning convergence and system utility compared with benchmark methods.

AIGC调度扩散模型能耗优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。