arXiv:2607.02941cs.AI2026-07

用滑动窗口强化学习解决多产品装配动态调度难题

A Sliding-Window-Based Reinforcement Learning for Dynamic Assembly Flow Shop Scheduling with Multi-Product Delivery

论文配图:A Sliding-Window-Based Reinforcement Learning for Dynamic Assembly Flow Shop Scheduling with Multi-Product Delivery
图 1 · 摘自论文原文
  • 基于滑动窗口过滤无效节点,聚焦关键装配操作
  • 在真实家电制造数据上实现延迟率显著降低
  • 适合需要实时响应的柔性装配流水线场景

多产品套件配送给集成加工与装配的混合制造系统带来实时调度挑战,因动态订单到达同时改变供应依赖关系和可行的任务-机器分配集。本文提出一种基于滑动窗口的强化学习(SWRL)框架,用于具有复杂套件约束的柔性装配流水线调度的端到端在线调度。该问题被建模为异构图基马尔可夫决策过程,捕捉双层套件结构及尾部产品瓶颈动态,导致稀疏奖励景观。为应对挑战,SWRL集成滑动窗口过滤机制,筛选非活跃节点并优先处理套件关键操作;时空图编码网络追踪连续决策状态中的瓶颈转移;动态动作映射模块结合受限等待策略,适应可变拓扑下的动作空间变化。在某家电制造商的真实案例上实验表明,SWRL在延迟率方面持续优于经典派工规则与现有深度强化学习方法,并在不同资源配置、订单负载和到达集中度下均表现出鲁棒性能。

原文摘要 · Abstract (English)

Multi-product kitting delivery imposes significant challenges for real-time scheduling in hybrid manufacturing systems that integrate processing and assembly, as dynamic order arrivals simultaneously alter supply dependencies and the set of feasible job-machine assignments. This paper proposes a sliding-window-based reinforcement learning (SWRL) framework for end-to-end online scheduling in the flexible assembly flow shop scheduling problem with complex kitting constraints. The problem is formulated as a heterogeneous graph-based Markov decision process that captures the dual-layer kitting structure and the tail-product bottleneck dynamics that produce a sparse reward landscape. To address the resulting challenges, SWRL integrates a sliding-window filtering mechanism that filters inactive nodes and prioritizes kitting-critical operations, a spatiotemporal graph encoding network that tracks bottleneck shifts across consecutive decision states, and a dynamic action mapping module with a constrained waiting strategy that adapts to the changing action space under variable topologies. Experiments on real-world instances from a home appliance manufacturer demonstrate that SWRL achieves consistent tardiness reductions over classical dispatching rules and existing deep reinforcement learning methods, and exhibits robust performance across varying resource configurations, order loads, and arrival concentrations.

动态调度强化学习装配流水线

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。