用持续迁移学习提升集群调度实时性,降低延迟。
Enhancing Cluster Scheduling in HPC: A Continuous Transfer Learning for Real-Time Optimization
- 基于持续迁移学习动态优化任务调度
- 在谷歌集群数据上达99%准确率,降低计算开销
- 适合需要实时调度的高性能计算场景
本研究提出一种机器学习辅助的集群任务调度优化方法,重点关注节点亲和性约束。传统调度器如Kubernetes在实时适应性方面表现不足,而所提出的持续迁移学习模型可在运行过程中动态演化,显著减少重新训练需求。在Google Cluster Data上的评估显示,该模型准确率超过99%,有效降低计算开销并改善受限任务的调度延迟。此可扩展方案支持实时优化,推动机器学习在集群管理中的集成,为未来自适应调度策略铺平道路。
原文摘要 · Abstract (English)
This study presents a machine learning-assisted approach to optimize task scheduling in cluster systems, focusing on node-affinity constraints. Traditional schedulers like Kubernetes struggle with real-time adaptability, whereas the proposed continuous transfer learning model evolves dynamically during operations, minimizing retraining needs. Evaluated on Google Cluster Data, the model achieves over 99% accuracy, reducing computational overhead and improving scheduling latency for constrained tasks. This scalable solution enables real-time optimization, advancing machine learning integration in cluster management and paving the way for future adaptive scheduling strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。