arXiv:2504.15021cs.DCcs.LG2025-04被引 1

用机器学习智能调度云服务器多资源,提升性能并降低开销

Is Intelligence the Right Direction in New OS Scheduling for Multiple Resources in Cloud Environments?

  • 采用多模型协同学习,同时调度缓存、内存和计算资源
  • 相比以往方法,收敛更快,支持更高负载且满足服务质量要求
  • 可动态适应工作负载变化,跨不同服务器部署效果好

针对云计算环境中多资源协同调度的挑战,本文提出OSML+,一种基于机器学习的新颖资源调度机制。该机制在内存层次结构中智能协调缓存与主内存带宽资源,并同步调度计算核心资源。通过多模型协同学习策略,OSML+能有效应对复杂场景,如避免资源瓶颈、应用间资源共享以及为不同优先级应用实施差异化调度策略。实验表明,相比先前研究,OSML+具备更快的收敛速度,可在更低开销下支持更高负载并达成服务质量目标。借助迁移学习技术,系统展现出良好的跨平台兼容性,适用于各类主流大规模云服务器。在真实负载下的测试验证了其在动态工作负载中的自适应能力与高效性。

原文摘要 · Abstract (English)

Making it intelligent is a promising way in System/OS design. This paper proposes OSML+, a new ML-based resource scheduling mechanism for co-located cloud services. OSML+ intelligently schedules the cache and main memory bandwidth resources at the memory hierarchy and the computing core resources simultaneously. OSML+ uses a multi-model collaborative learning approach during its scheduling and thus can handle complicated cases, e.g., avoiding resource cliffs, sharing resources among applications, enabling different scheduling policies for applications with different priorities, etc. OSML+ can converge faster using ML models than previous studies. Moreover, OSML+ can automatically learn on the fly and handle dynamically changing workloads accordingly. Using transfer learning technologies, we show our design can work well across various cloud servers, including the latest off-the-shelf large-scale servers. Our experimental results show that OSML+ supports higher loads and meets QoS targets with lower overheads than previous studies.

资源调度机器学习云计算系统优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。