用多智能体强化学习优化机器调度,提升复杂场景下的可扩展性。
Exploring Multi-Agent Reinforcement Learning for Unrelated Parallel Machine Scheduling
- 采用多智能体强化学习框架建模机器调度问题
- 多智能体方法在复杂场景下展现更强可扩展性
- 适合需要智能调度的工业制造与资源管理场景
调度问题在资源、工业和运营管理中具有重要挑战。本文针对带有准备时间与资源约束的非相关并行机调度问题(UPMS),采用多智能体强化学习(MARL)方法进行求解。研究构建了强化学习环境,并通过实验对比了多智能体与单智能体算法的表现。实验使用多种深度神经网络策略,分别应用于单智能体与多智能体方法。结果表明,在单智能体场景中,可遮蔽的近端策略优化(Maskable PPO)算法表现优异;在多智能体设置中,多智能体PPO算法展现出良好性能。尽管单智能体算法在简化场景下表现尚可,但多智能体方法在协作学习上存在挑战,同时具备更强的可扩展性。本研究为将MARL应用于调度优化提供了新视角,强调算法复杂度与可扩展性之间的平衡对实现智能化调度至关重要。
原文摘要 · Abstract (English)
Scheduling problems pose significant challenges in resource, industry, and operational management. This paper addresses the Unrelated Parallel Machine Scheduling Problem (UPMS) with setup times and resources using a Multi-Agent Reinforcement Learning (MARL) approach. The study introduces the Reinforcement Learning environment and conducts empirical analyses, comparing MARL with Single-Agent algorithms. The experiments employ various deep neural network policies for single- and Multi-Agent approaches. Results demonstrate the efficacy of the Maskable extension of the Proximal Policy Optimization (PPO) algorithm in Single-Agent scenarios and the Multi-Agent PPO algorithm in Multi-Agent setups. While Single-Agent algorithms perform adequately in reduced scenarios, Multi-Agent approaches reveal challenges in cooperative learning but a scalable capacity. This research contributes insights into applying MARL techniques to scheduling optimization, emphasizing the need for algorithmic sophistication balanced with scalability for intelligent scheduling solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。