用多智能体强化学习动态调度云原生数据库资源,提升效率与稳定性。
Multi-Agent Reinforcement Learning for Adaptive Resource Orchestration in Cloud-Native Clusters
- 设计异构角色智能体模型,让不同资源节点采用不同策略。
- 结合局部性能与全局反馈优化奖励,加快策略收敛并减少偏差。
- 在真实生产数据上验证,适用于高并发复杂依赖场景。
本文针对云原生数据库系统中资源高度动态和调度复杂的问题,提出一种基于多智能体强化学习的自适应资源编排方法。该方法引入异构角色驱动的智能体建模机制,使计算节点、存储节点与调度器等不同资源实体能够采用差异化的策略表示,更准确反映其功能职责与局部环境特征。设计了融合局部观测与全局反馈的奖励重塑机制,缓解因状态观测不全导致的策略学习偏差。通过实时局部性能信号与全局系统价值估计的结合,增强智能体间协调性,提升策略收敛稳定性。构建统一的多智能体训练框架,在典型生产调度数据集上进行评估。实验表明,所提方法在资源利用率、调度延迟、策略收敛速度、系统稳定性和公平性等多个关键指标上均优于传统方法。结果证实其具备强泛化能力与实用价值,在高并发、高维状态空间及复杂依赖关系的多种场景下表现优异,适用于真实大规模调度环境。
原文摘要 · Abstract (English)
This paper addresses the challenges of high resource dynamism and scheduling complexity in cloud-native database systems. It proposes an adaptive resource orchestration method based on multi-agent reinforcement learning. The method introduces a heterogeneous role-based agent modeling mechanism. This allows different resource entities, such as compute nodes, storage nodes, and schedulers, to adopt distinct policy representations. These agents are better able to reflect diverse functional responsibilities and local environmental characteristics within the system. A reward-shaping mechanism is designed to integrate local observations with global feedback. This helps mitigate policy learning bias caused by incomplete state observations. By combining real-time local performance signals with global system value estimation, the mechanism improves coordination among agents and enhances policy convergence stability. A unified multi-agent training framework is developed and evaluated on a representative production scheduling dataset. Experimental results show that the proposed method outperforms traditional approaches across multiple key metrics. These include resource utilization, scheduling latency, policy convergence speed, system stability, and fairness. The results demonstrate strong generalization and practical utility. Across various experimental scenarios, the method proves effective in handling orchestration tasks with high concurrency, high-dimensional state spaces, and complex dependency relationships. This confirms its advantages in real-world, large-scale scheduling environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。