对比联合与模块化训练,揭示调度最优的适用条件。
An Analysis of the Coordination Gap between Joint and Modular Learning for Job Shop Scheduling with Transportation Resources
- 通过资源稀缺性与时间主导性的敏感性分析,量化两种训练方式差异
- 联合训练在多数场景下优于模块化训练,但瓶颈环境下优势减弱
- 为制造调度中训练策略选择提供实证依据,适合工业部署决策者
高效处理带运输资源的作业车间调度对高性能制造至关重要。随着‘去中心化工厂’兴起,多智能体强化学习成为协同生产与运输调度的有前景方法。以往研究多聚焦于新型协作架构,忽视了联合训练是否必要这一关键问题。联合训练指同时训练工件与自动导引车调度智能体,而模块化训练则是分别独立训练后后期集成。本文系统研究在何种条件下联合训练对作业车间调度性能最优。通过资源稀缺性和时间主导性严格的敏感性分析,量化了两种训练模式间的协调差距——性能差异。实验表明,联合训练优于多数调度规则组合及模块化训练方法;但在瓶颈环境,尤其是严重运输与加工约束下,协调差距优势减弱。这表明当单一调度任务占主导时,模块化训练是可行替代方案。本研究为基于强化学习的调度策略选择提供了实践指导,可根据环境条件优化性能。
原文摘要 · Abstract (English)
Efficient job-shop scheduling with transportation resources is critical for high-performance manufacturing. With the rise of "decentralized factories", multi-agent reinforcement learning has emerged as a promising approach for the combined scheduling of production and transportation tasks. Prior work has largely focused on developing novel cooperative architectures while overlooking the question of when joint training is necessary. Joint training denotes the simultaneous training of job and automatic guided vehicle scheduling agents, whereas modular training involves independently training each agent followed by post-hoc integration. In this study, we systematically investigate the conditions under which joint training is essential for optimal performance in the job-shop scheduling problem with transportation resources. Through a rigorous sensitivity analysis of resource scarcity and temporal dominance, we quantify the coordination gap -- the performance difference between these two training modalities. In our evaluation, joint training outperforms the majority of dispatching rule combinations and modular training approaches. However, the coordination gap advantage diminishes in bottleneck environments, particularly under severe transport and processing constraints. These findings indicate that modular training represents a viable alternative in environments where a single scheduling task dominates. Overall, our work provides practical guidance for selecting between training modalities based on environmental conditions, enabling decision-makers to optimize reinforcement learning-based scheduling performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。