arXiv:2511.07071cs.MAcs.RO2025-11被引 1

用多智能体强化学习解决自动导引车系统的死锁问题,提升物流系统灵活性与效率。

Multi-Agent Reinforcement Learning for Deadlock Handling among Autonomous Mobile Robots

  • 采用MARL结合集中训练分散执行策略,动态应对复杂拥堵场景
  • 在密集环境中,MARL方案比规则方法吞吐量提升30%以上
  • 适合高动态、高密度的智能物流系统部署,尤其适用于复杂路径规划

本论文研究多智能体强化学习(MARL)在依赖自主移动机器人(AMR)的物流系统中处理死锁的应用。AMR虽提升运营灵活性,却增加死锁风险,降低系统吞吐量与可靠性。现有方法常忽视规划阶段的死锁处理,依赖僵化控制规则,难以适应动态工况。为此,本文提出一套将MARL融入物流规划与控制的系统化方法,构建显式包含死锁能力的多智能体路径规划(MAPF)参考模型,支持对MARL策略的系统评估。基于网格环境与外部仿真软件,对比传统死锁处理方式与MARL方案,重点分析PPO与IMPALA算法在不同训练/执行模式下的表现。结果表明,在复杂拥挤环境下,尤其是采用集中训练分散执行(CTDE)的MARL策略显著优于规则方法;而在简单或空间充裕环境中,规则方法因计算开销低仍具竞争力。研究证明,MARL为动态物流场景中的死锁处理提供了灵活可扩展的解决方案,但需根据实际运行条件进行适配。

原文摘要 · Abstract (English)

This dissertation explores the application of multi-agent reinforcement learning (MARL) for handling deadlocks in intralogistics systems that rely on autonomous mobile robots (AMRs). AMRs enhance operational flexibility but also increase the risk of deadlocks, which degrade system throughput and reliability. Existing approaches often neglect deadlock handling in the planning phase and rely on rigid control rules that cannot adapt to dynamic operational conditions. To address these shortcomings, this work develops a structured methodology for integrating MARL into logistics planning and operational control. It introduces reference models that explicitly consider deadlock-capable multi-agent pathfinding (MAPF) problems, enabling systematic evaluation of MARL strategies. Using grid-based environments and an external simulation software, the study compares traditional deadlock handling strategies with MARL-based solutions, focusing on PPO and IMPALA algorithms under different training and execution modes. Findings reveal that MARL-based strategies, particularly when combined with centralized training and decentralized execution (CTDE), outperform rule-based methods in complex, congested environments. In simpler environments or those with ample spatial freedom, rule-based methods remain competitive due to their lower computational demands. These results highlight that MARL provides a flexible and scalable solution for deadlock handling in dynamic intralogistics scenarios, but requires careful tailoring to the operational context.

多智能体强化学习物流优化死锁处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。