arXiv:2503.00537cs.LG2025-03

提出可扩展的强化学习框架,解决大规模虚拟机调度难题。

Scalable Reinforcement Learning for Virtual Machine Scheduling

  • 采用分解+前瞻+Top-k筛选三重机制应对大规模调度复杂性
  • 支持最多50台物理机,较现有方法提升4倍规模上限
  • 在多种场景下超越当前最优方法,适合云平台资源调度研究者

近期强化学习(RL)在小规模集群的虚拟机调度(VMS)中展现出潜力,但在大规模云计算场景中的应用仍受限。本文提出一种可扩展的强化学习框架——集群价值分解强化学习(CVD-RL),以突破大规模VMS的可扩展性瓶颈。该框架创新性地结合分解算子与前瞻算子,有效管理表示复杂性,同时辅以Top-$k$过滤算子提升探索效率。不同于以往仅限于10台或以下物理机(PMs)集群的方法,CVD-RL可扩展至最多50台PM的环境。实证研究表明,该框架在多种场景下均表现出优于当前SOTA方法的泛化能力。这一突破不仅验证了框架在复杂大规模云基础设施中卓越的可扩展性与性能,更标志着强化学习在大规模虚拟机调度应用上的重要进展。代码已公开于https://anonymous.4open.science/r/marl4sche-D0FE。

原文摘要 · Abstract (English)

Recent advancements in reinforcement learning (RL) have shown promise for optimizing virtual machine scheduling (VMS) in small-scale clusters. The utilization of RL to large-scale cloud computing scenarios remains notably constrained. This paper introduces a scalable RL framework, called Cluster Value Decomposition Reinforcement Learning (CVD-RL), to surmount the scalability hurdles inherent in large-scale VMS. The CVD-RL framework innovatively combines a decomposition operator with a look-ahead operator to adeptly manage representation complexities, while complemented by a Top-$k$ filter operator that refines exploration efficiency. Different from existing approaches limited to clusters of $10$ or fewer physical machines (PMs), CVD-RL extends its applicability to environments encompassing up to $50$ PMs. Furthermore, the CVD-RL framework demonstrates generalization capabilities that surpass contemporary SOTA methodologies across a variety of scenarios in empirical studies. This breakthrough not only showcases the framework's exceptional scalability and performance but also represents a significant leap in the application of RL for VMS within complex, large-scale cloud infrastructures. The code is available at https://anonymous.4open.science/r/marl4sche-D0FE.

强化学习虚拟机调度云平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。