解决模块化工厂中因养护等导致的延迟对调度的干扰,提升排产效率。
Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Module Factories
- 引入滞后感知机制与双注意力强化学习,显式建模生产延迟影响
- 在基准测试中比传统规则快4%接近最优解,且在资源紧张时优势更明显
- 无需依赖求解器或许可证,可秒级重规划,适合真实工业场景
预制装配式建筑将大量施工环节转移至模块工厂,其生产流程呈现柔性作业车间特性。主要挑战在于混凝土养护、防水试验、油漆干燥等造成的长后处理延迟,此时模块被占用而工作站空闲。基于国家预制规范的基准实例显示,这些延迟使最优调度周期平均增加约67%,忽略延迟直接决策再修复可行性的做法,反而劣于所有经典调度规则。本文通过三项轻量级可独立验证的改进,对前沿双注意力深度强化学习求解器进行适配:滞后感知动态与可接受奖励边界、两个前瞻滞后特征通道、以及操作与工位类型掩码嵌入。每项改动均可关闭以复现原求解器,所有性能提升均归因于新增设计。我们公开了基于规范的基准生成工具。在未见实例上,该学习策略成为无需求解器的最强调度器:达到约束规划参考解的约96%,优于所有调度规则与遗传算法元启发式,且在高负载下优势扩大;单一混合规模策略可覆盖训练范围内的全部工厂规模。系统无需求解器、模型或许可证在环运行,突发扰动后可在数秒内完成重规划;当精确求解器可用时,其仍为质量上限,本文明确绘制了这一边界。
原文摘要 · Abstract (English)
Prefabricated prefinished volumetric construction moves most building work into module factories, whose production floor operates as a flexible job shop. A major complication is decisive: long post-operation time-lags caused by concrete curing, watertightness ponding tests, and paint drying, during which a module is blocked while its workstation stays free. On benchmark instances grounded in an official national prefabrication guidebook, these lags inflate even the optimal reference makespan by about 67% on average, and ignoring them at decision time, then repairing to feasibility, is worse than every dispatching rule. We adapt a state-of-the-art dual-attention deep reinforcement learning solver through three minimally invasive, individually ablatable extensions: lag-aware dynamics with an admissible reward bound, two anticipatory lag feature channels, and liveness-masked operation- and station-type embeddings. With every extension disabled the implementation reproduces the original solver exactly, so all gains are attributable to the adaptations. We release a public, guidebook-grounded benchmark generator. On held-out instances the learned policy is the strongest solver-free scheduler: it reaches within about 4% of a constraint-programming reference and beats every dispatching rule and a genetic-algorithm metaheuristic, with its advantage widening under capacity contention, and a single size-mixed policy carries this lead across the trained range of factory sizes. It needs no solver, model, or license in the loop and re-plans within seconds of a disruption; where an exact solver can be deployed, that solver remains the quality ceiling, a boundary we map explicitly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。