arXiv:2604.09202cs.LGcs.AI2026-04

用图神经网络调度云任务时,结构差异会导致失败。

On the Role of DAG topology in Energy-Aware Cloud Scheduling : A GNN-Based Deep Reinforcement Learning Approach

  • 用GNN+强化学习自动调度云工作流
  • 发现训练与部署结构不同时性能严重下降
  • 适合研究云调度鲁棒性的工程师

云服务提供商需在满足完成时间、成本和能耗等多重目标的前提下,将异构计算资源分配给工作流有向无环图(DAG)。本文研究单个工作流、无队列的调度场景,提出一种基于图神经网络(GNN)的深度强化学习调度器,旨在最小化工作流完成时间和能耗。通过受控的分布外(OOD)评估,我们识别出该类调度器在特定分布外条件下会失效,并从理论上解释了失败原因:训练与部署环境间的结构差异破坏了消息传递机制,导致策略泛化能力下降。分析揭示了当前GNN调度器的根本局限性,强调需发展更鲁棒的表示方法以应对分布偏移,保障调度可靠性。

原文摘要 · Abstract (English)

Cloud providers must assign heterogeneous compute resources to workflow DAGs while balancing competing objectives such as completion time, cost, and energy consumption. In this work, we study a single-workflow, queue-free scheduling setting and consider a graph neural network (GNN)-based deep reinforcement learning scheduler designed to minimize workflow completion time and energy usage. We identify specific out-of-distribution (OOD) conditions under which GNN-based deep reinforcement learning schedulers fail and provide a principled explanation of why these failures occur. Through controlled OOD evaluations, we demonstrate that performance degradation stems from structural mismatches between training and deployment environments, which disrupt message passing and undermine policy generalization. Our analysis exposes fundamental limitations of current GNN-based schedulers and highlights the need for more robust representations to ensure reliable scheduling performance under distribution shifts.

云调度GNN强化学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。