arXiv:2412.19538cs.ROcs.AI2024-12被引 2

提出分层强化学习框架,解决超大规模机器人仓储系统任务规划难题。

Scalable Hierarchical Reinforcement Learning for Hyper Scale Multi-Robot Task Planning

  • 构建分层时序注意力网络与多阶段课程学习提升扩展性
  • 在200机器人1000货架的未学地图上仍保持优异性能
  • 适合大规模仓储机器人协同调度场景

为提升仓储系统效率并应对海量客户订单,本文针对机器人移动履约系统(RMFS)中超大规模多机器人任务规划(MRTP)面临的维度灾难与动态性挑战,提出一种基于分层强化学习(HRL)的高效多阶段任务规划方法。规划过程采用特殊的时间图拓扑结构表示。设计集中式架构以保证最优性,但面临可扩展性与泛化性挑战,需在未见过的规模和地图上维持性能。为此,首先构建分层时序注意力网络(HTAN)以处理不定长输入;其次设计多阶段课程学习策略,增强扩展能力并避免灾难性遗忘。此外,针对分层策略中存在的信用分配不公问题,提出基于反事实回放基线的分层强化学习算法,提升学习效率。实验表明,该规划器在模拟与真实RMFS环境下均优于现有先进方法,且可在未学地图上成功扩展至200机器人、1000个取货架的超大规模实例,性能显著领先。

原文摘要 · Abstract (English)

To improve the efficiency of warehousing system and meet huge customer orders, we aim to solve the challenges of dimension disaster and dynamic properties in hyper scale multi-robot task planning (MRTP) for robotic mobile fulfillment system (RMFS). Existing research indicates that hierarchical reinforcement learning (HRL) is an effective method to reduce these challenges. Based on that, we construct an efficient multi-stage HRL-based multi-robot task planner for hyper scale MRTP in RMFS, and the planning process is represented with a special temporal graph topology. To ensure optimality, the planner is designed with a centralized architecture, but it also brings the challenges of scaling up and generalization that require policies to maintain performance for various unlearned scales and maps. To tackle these difficulties, we first construct a hierarchical temporal attention network (HTAN) to ensure basic ability of handling inputs with unfixed lengths, and then design multi-stage curricula for hierarchical policy learning to further improve the scaling up and generalization ability while avoiding catastrophic forgetting. Additionally, we notice that policies with hierarchical structure suffer from unfair credit assignment that is similar to that in multi-agent reinforcement learning, inspired of which, we propose a hierarchical reinforcement learning algorithm with counterfactual rollout baseline to improve learning performance. Experimental results demonstrate that our planner outperform other state-of-the-art methods on various MRTP instances in both simulated and real-world RMFS. Also, our planner can successfully scale up to hyper scale MRTP instances in RMFS with up to 200 robots and 1000 retrieval racks on unlearned maps while keeping superior performance over other methods.

多机器人强化学习仓储系统规模化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。