用多智能体强化学习实现工厂级动态调度,提升规模与鲁棒性。
Scalable Multi-agent Reinforcement Learning for Factory-wide Dynamic Scheduling
- 分解任务为子问题,通过领导-跟随机制实现多智能体协同
- 在多种场景下优于现有深度强化学习调度模型
- 对需求变化具有强鲁棒性,适合实际制造场景
实时动态调度是现代制造过程中的关键但极具挑战性的任务,因其决策复杂度高而难以处理。近年来,强化学习(RL)因其在应对该挑战方面的潜力受到关注。然而,传统RL方法通常依赖人工制定的调度规则,难以适用于大规模工厂级调度。为此,本文采用领导-跟随式多智能体强化学习(MARL)框架,将调度问题分解为多个子问题,由各智能体独立处理,以实现可扩展性。为进一步增强系统稳定性,提出基于规则的转换算法,防止因单个智能体失误导致生产产能大幅下降。实验结果表明,所提模型在多个方面均优于当前最先进的基于深度强化学习的调度方法,且对需求波动表现出最强的鲁棒性。整体而言,该基于MARL的调度模型为实时调度问题提供了有前景的解决方案,具备在多种制造业中的应用潜力。
原文摘要 · Abstract (English)
Real-time dynamic scheduling is a crucial but notoriously challenging task in modern manufacturing processes due to its high decision complexity. Recently, reinforcement learning (RL) has been gaining attention as an impactful technique to handle this challenge. However, classical RL methods typically rely on human-made dispatching rules, which are not suitable for large-scale factory-wide scheduling. To bridge this gap, this paper applies a leader-follower multi-agent RL (MARL) concept to obtain desired coordination after decomposing the scheduling problem into a set of sub-problems that are handled by each individual agent for scalability. We further strengthen the procedure by proposing a rule-based conversion algorithm to prevent catastrophic loss of production capacity due to an agent's error. Our experimental results demonstrate that the proposed model outperforms the state-of-the-art deep RL-based scheduling models in various aspects. Additionally, the proposed model provides the most robust scheduling performance to demand changes. Overall, the proposed MARL-based scheduling model presents a promising solution to the real-time scheduling problem, with potential applications in various manufacturing industries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。