用图神经网络和强化学习优化多目标排产,兼顾准时率与换产时间。
Graph-Enhanced Deep Reinforcement Learning for Multi-Objective Unrelated Parallel Machine Scheduling
- 用图神经网络建模工件、机器与换产关系,强化学习直接学排产策略。
- 在基准测试中同时降低总加权延误和总换产时间,优于传统方法。
- 适合复杂制造调度场景,可扩展性强,无需手动设计规则。
带释放时间、换产和准入约束的非同质并行机调度问题(UPMSP)是典型的多目标难题。传统方法难以平衡总加权延误(TWT)与总换产时间(TST)。本文提出基于近端策略优化(PPO)与图神经网络(GNN)的深度强化学习框架。GNN有效表征工件、机器与换产的复杂状态,使PPO智能体直接学习调度策略。在多目标奖励函数引导下,智能体同时最小化TWT与TST。在基准实例上的实验表明,该PPO-GNN代理显著优于标准派工规则和元启发式算法,在两个目标间实现更优权衡,为复杂制造调度提供鲁棒且可扩展的解决方案。
原文摘要 · Abstract (English)
The Unrelated Parallel Machine Scheduling Problem (UPMSP) with release dates, setups, and eligibility constraints presents a significant multi-objective challenge. Traditional methods struggle to balance minimizing Total Weighted Tardiness (TWT) and Total Setup Time (TST). This paper proposes a Deep Reinforcement Learning framework using Proximal Policy Optimization (PPO) and a Graph Neural Network (GNN). The GNN effectively represents the complex state of jobs, machines, and setups, allowing the PPO agent to learn a direct scheduling policy. Guided by a multi-objective reward function, the agent simultaneously minimizes TWT and TST. Experimental results on benchmark instances demonstrate that our PPO-GNN agent significantly outperforms a standard dispatching rule and a metaheuristic, achieving a superior trade-off between both objectives. This provides a robust and scalable solution for complex manufacturing scheduling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。