用多智能体强化学习实现可重构产线的动态调度,实时应对设备故障和重构延迟。
A Negotiation-Based Multi-Agent Reinforcement Learning Approach for Dynamic Scheduling of Reconfigurable Manufacturing Systems
- 设计基于注意力机制的谈判框架,提升智能体对关键系统特征的决策聚焦。
- 在模拟环境中使完工时间缩短31%,延迟率降低27%,机器利用率提升19%。
- 适合研究智能制造调度、强化学习应用或工业自动化系统的读者。
可重构制造系统(RMS)因其能快速响应市场需求变化、新技术引入及供应链中断而成为未来制造的关键。其可调硬件需配合灵活软件规划机制,以实现实时生产计划与调度。本文探索多智能体强化学习(MARL)在RMS软规划中的动态调度应用。提出框架中,采用集中训练的深度Q网络(DQN)智能体实时学习最优任务-机器分配,适应如设备故障和重构延迟等随机事件。模型引入带注意力机制的谈判机制,增强状态表征并聚焦关键系统特征。关键改进包括优先经验回放、n步回报、双DQN和软目标更新,以稳定并加速学习。在模拟RMS环境中的实验表明,该方法相比基线启发式算法显著降低完工时间(减少31%)和延迟率(降低27%),同时提升机器利用率(提高19%)。扩展环境模拟了真实挑战,如设备故障和重构时间。结果表明,尽管增强后的DQN智能体能有效应对动态条件,但设备故障仍会增加完工时间、吞吐量和总延迟的波动性。研究证实了将MARL应用于动态可重构制造环境智能自适应调度的优势。
原文摘要 · Abstract (English)
Reconfigurable manufacturing systems (RMS) are critical for future market adjustment given their rapid adaptation to fluctuations in consumer demands, the introduction of new technological advances, and disruptions in linked supply chain sections. The adjustable hard settings of such systems require a flexible soft planning mechanism that enables realtime production planning and scheduling amid the existing complexity and variability in their configuration settings. This study explores the application of multi agent reinforcement learning (MARL) for dynamic scheduling in soft planning of the RMS settings. In the proposed framework, deep Qnetwork (DQN) agents trained in centralized training learn optimal job machine assignments in real time while adapting to stochastic events such as machine breakdowns and reconfiguration delays. The model also incorporates a negotiation with an attention mechanism to enhance state representation and improve decision focus on critical system features. Key DQN enhancements including prioritized experience replay, nstep returns, double DQN and soft target update are used to stabilize and accelerate learning. Experiments conducted in a simulated RMS environment demonstrate that the proposed approach outperforms baseline heuristics in reducing makespan and tardiness while improving machine utilization. The reconfigurable manufacturing environment was extended to simulate realistic challenges, including machine failures and reconfiguration times. Experimental results show that while the enhanced DQN agent is effective in adapting to dynamic conditions, machine breakdowns increase variability in key performance metrics such as makespan, throughput, and total tardiness. The results confirm the advantages of applying the MARL mechanism for intelligent and adaptive scheduling in dynamic reconfigurable manufacturing environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。