用真实涂装场景验证强化学习调度,提升工业适用性。
Reinforcement Learning-Based Production Scheduling in an Industry-Based Coating Scenario Using the Digital Model Playground

- 基于数字沙盒仿真训练强化学习模型,模拟真实生产复杂性。
- PPO算法在多指标上表现最优,整体性能优于传统规则。
- 开源框架可复现,适合工业界与学术界协同研究。
复杂制造环境中的生产调度面临工序依赖的准备时间、随机干扰和交期约束等多重挑战。尽管强化学习(RL)在研究中展现出潜力,但多数工作依赖简化基准流程,缺乏工业相关性。本文在贴近实际的涂装工艺场景中验证了基于强化学习的调度方法,该场景包含工序依赖的准备时间、设备故障和利用率波动等现实复杂性。采用开源的数字模型沙盒(DMPG)作为离散事件仿真框架,对深度Q网络(DQN)和近端策略优化(PPO)两种标准算法进行训练,并与传统调度规则对比,验证其可行性并提供透明可复现的测试平台。结果表明,强化学习调度在关键绩效指标上实现均衡提升,其中PPO表现最稳健。本工作的主要贡献在于连接学术研究与工业实践,通过在真实可复现场景中验证强化学习调度,并提供可重用的开源框架,推动后续研究发展。
原文摘要 · Abstract (English)
Production scheduling in complex manufacturing environments is challenging when sequence-dependent setup times, stochastic disturbances, and due-date constraints must be addressed simultaneously. While reinforcement learning (RL) methods have shown promising results in research, most studies rely on simplified benchmark processes, limiting their industrial relevance. This paper demonstrates the applicability of RL-based scheduling in an industry-inspired coating process that reflects practical complexities such as sequence-dependent setup times, machine breakdowns, and variable utilization. The open-source Digital Model Playground (DMPG), a discrete event simulation framework, is used to model the scenario and to train RL agents. Two standard algorithms, Deep Q-Networks and Proximal Policy Optimization, are benchmarked against conventional dispatching rules to illustrate feasibility and to provide a transparent testbed for further research. Results indicate that RL-based scheduling achieves balanced improvements across key performance indicators, with PPO delivering the most robust performance. The main contribution of this work is to bridge the gap between academic research and industrial practice by validating RL-based scheduling in a realistic, shareable scenario and by providing a reusable open-source framework for future studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。