用离散事件模拟构建高效强化学习环境,提升车间调度求解速度与精度。
An efficient deep reinforcement learning environment for flexible job-shop scheduling
- 基于离散事件仿真构建时序化调度环境,简化状态建模。
- 在公开数据集上性能超越传统优先规则,媲美最优算法。
- 适合需快速生成高质量调度方案的工业场景应用。
柔性作业车间调度问题(FJSP)是经典的组合优化问题,广泛应用于实际生产中。为快速准确生成调度方案,已有多种深度强化学习(DRL)调度方法被提出。然而,这些方法多聚焦于DRL智能体设计,忽视了对DRL环境的建模。本文基于离散事件仿真提出一种简洁的时序化DRL环境,并构建端到端的调度模型,采用近端策略优化(PPO)算法。同时,提出一种基于两个状态变量的新状态表示,以及基于机器作业区域的可解释奖励函数。在公开基准实例上的实验表明,该调度环境提升了简单优先规则(PDR)的表现,所提DRL模型性能与OR-Tools、元启发式、其他DRL及PDR方法相当。
原文摘要 · Abstract (English)
The Flexible Job-shop Scheduling Problem (FJSP) is a classical combinatorial optimization problem that has a wide-range of applications in the real world. In order to generate fast and accurate scheduling solutions for FJSP, various deep reinforcement learning (DRL) scheduling methods have been developed. However, these methods are mainly focused on the design of DRL scheduling Agent, overlooking the modeling of DRL environment. This paper presents a simple chronological DRL environment for FJSP based on discrete event simulation and an end-to-end DRL scheduling model is proposed based on the proximal policy optimization (PPO). Furthermore, a short novel state representation of FJSP is proposed based on two state variables in the scheduling environment and a novel comprehensible reward function is designed based on the scheduling area of machines. Experimental results on public benchmark instances show that the performance of simple priority dispatching rules (PDR) is improved in our scheduling environment and our DRL scheduling model obtains competing performance compared with OR-Tools, meta-heuristic, DRL and PDR scheduling methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。