arXiv:2608.30512cs.LGcs.AI2026-08

用神经网络统一调度路径,提升大规模吊车系统的运输效率。

Trajectory-Initialized Neural Double Q-Routing for Large-Scale Overhead Hoist Transport Systems

论文配图:Trajectory-Initialized Neural Double Q-Routing for Large-Scale Overhead Hoist Transport Systems
图 1 · 摘自论文原文
  • 用共享神经网络替代传统表格,实现全局状态价值学习。
  • 在150至200台吊车场景中,平均完成时间降低0.8%至8.8%。
  • 支持快速启动,任务完成数接近原方法,尾部延迟显著减少。

大规模工业机器人车队共享有限物理基础设施,导致车辆行驶时间受安全间隔、交叉口访问、下游阻塞和站点争用的影响。本文研究半导体晶圆厂中典型的悬挂式吊车运输(OHT)系统。静态最短路径路由无法反映这些时变交通成本,而表格型Q路由虽可在线适应,但每个目的地-节点-动作值独立学习,限制了稀疏路径间的知识共享,且初始行为对不准确估值敏感。本文提出神经双Q路由(Neural Double Q-routing),将目的地索引表替换为共享的状态-动作价值网络。该网络通过混合模拟器生成的路径进行回溯值回归进行预热,再结合双Q更新、局部拥塞修正与事件分层结构回放在线优化。在九组匹配的车队规模-到达率设置(100、150、200台OHT)下,相比表格型双Q路由,平均完成时间降低0.8%–8.8%。在六个150及200台设置中,该方法取得最低平均完成时间;而100台设置中Dijkstra仍最优。任务完成数在八组设置中与表格型保持±1%内,第95百分位完成时间在八组中下降。在两个匹配的启动场景中,离线初始化使完成任务数最多提升23%,尾部完成时间最多缩短15%。

原文摘要 · Abstract (English)

Large-scale industrial robot fleets share constrained physical infrastructure, making vehicle travel times dependent on safety separation, intersection access, downstream blocking, and station contention. We study this problem in overhead hoist transport (OHT) systems, a representative ceiling-mounted material-handling system used in semiconductor fabs. Static shortest-path routing cannot account for these time-varying traffic costs, whereas tabular Q-routing adapts online but learns each destination--node--action value independently, limiting information sharing across sparsely visited routing contexts and making startup behavior sensitive to inaccurate value estimates. We propose Neural Double Q-routing, which replaces destination-indexed tables with a shared state--action value network. The network is warm-started through return-to-go regression on mixed simulator-generated routing trajectories and then refined online using Double-Q updates, local congestion correction, and event-stratified structured replay. Across nine matched fleet-size--arrival-rate settings with 100, 150, and 200 OHTs, the proposed framework reduces mean completion time relative to tabular Double Q-routing by $0.8\%$--$8.8\%$. It achieves the lowest mean completion time among all compared methods in the six 150- and 200-OHT settings, whereas Dijkstra remains best in the three 100-OHT settings. Completed-task counts remain within $1\%$ of tabular Double Q-routing in eight of nine settings, and 95th-percentile completion time decreases in eight settings. In two matched startup scenarios, offline initialization increases the number of completed tasks by up to $23\%$ and reduces tail completion time by up to $15\%$.

智能调度强化学习工业系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。