arXiv:2608.05588cs.ROcs.AI2026-08

提出新方法解决机器人在仓库中持续移动时的避障与旋转难题。

Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations

论文配图:Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations
图 1 · 摘自论文原文
  • 用搜索增强神经策略,实时化解碰撞并传递意图
  • 联合优化机器人与环境策略,在高密度地图上成功率超基线37%
  • 适合需要长期运行的多机器人系统,如智能仓储

长期多机器人路径规划(LMAPF)需持续为不断接收新目标的机器人规划无碰撞路径。现有学习方法多基于简化运动假设,忽略真实场景中的运动约束。本文研究一种更贴近现实的模型LMAPF-R2,包含安全约束和原地旋转约束,显著提升协调难度,尤其在高密度空间中。为此提出搜索辅助联合强化学习(SJRL):首先用因果PIBT(Causal PIBT)作为单步搜索规划器,处理碰撞并传播意图;再引入统一强化学习框架,联合优化代理与环境策略,其中环境策略通过反向Dijkstra搜索学习图边代价,提供全局引导。实验表明,SJRL在多个高密度地图上显著优于强基线方法Causal-PIBT。进一步在混合现实仓库环境中验证,部署8台物理机器人与248台虚拟机器人,效果稳定可靠。

原文摘要 · Abstract (English)

Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals upon reaching their current ones. While many learning-based planners have been proposed for LMAPF, most rely on oversimplified kinematic assumptions that may overlook motion constraints critical to real-world performance. In this work, we study a more realistic LMAPF model derived from many real-world automated warehouse systems, termed LMAPF-R2, which incorporates robust safety constraints and in-place rotation constraints. These constraints substantially increase coordination difficulty, particularly in highly constrained spaces. To address these challenges, we propose Search-Aided Joint Reinforcement Learning (SJRL). We first augment neural policies with Causal PIBT, a single-step search-based planner that resolves agents' collisions and propagates their intentions. We then introduce a unified RL formulation that jointly optimizes agent and environment policies, where the environment policy learns graph edge costs to provide global movement guidance via backward Dijkstra search. Experiments demonstrate that SJRL achieves significant improvements over the strong search-based planner, Causal-PIBT, across multiple high-density maps. We further validate SJRL in a challenging mixed-reality warehouse environment with 8 physical robots and 248 virtual robots.

多智能体路径规划强化学习仓储机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。