arXiv:2601.03301cs.MAcs.AI2026-01中稿 · IROS 2025

提出PC2P框架,提升多智能体路径规划的通信与感知能力。

PC2P: Multi-Agent Path Finding via Personalized-Enhanced Communication and Crowd Perception

  • 基于动态图拓扑的个性化通信机制,明确交互中的'谁'与'什么'
  • 融合静态约束与动态占用的局部人群感知,增强决策引导
  • 区域化死锁突破策略,有效解决狭窄空间协同难题

分布式多智能体路径规划(MAPF)结合多智能体强化学习(MARL)已成为研究热点,可在部分可观测环境中通过智能体间通信实现实时协作决策。然而,现有方法因协作与感知能力不足,难以适应多样化环境。为此,本文提出基于Q-learning的新型分布式MAPF方法PC2P。首先,设计基于动态图拓扑的个性化增强通信机制,通过选择、生成、聚合三阶段操作,明确交互中的'谁'与'什么';同时引入局部人群感知,融合静态空间约束与动态占用变化,增强模型对有效动作的引导。为解决极端死锁问题,提出基于区域的死锁突破策略,利用专家指导在受限区域内实现高效协调。实验表明,PC2P在多种环境下的性能优于当前最优分布式MAPF方法。消融实验进一步验证了各模块对整体性能的有效性。

原文摘要 · Abstract (English)

Distributed Multi-Agent Path Finding (MAPF) integrated with Multi-Agent Reinforcement Learning (MARL) has emerged as a prominent research focus, enabling real-time cooperative decision-making in partially observable environments through inter-agent communication. However, due to insufficient collaborative and perceptual capabilities, existing methods are inadequate for scaling across diverse environmental conditions. To address these challenges, we propose PC2P, a novel distributed MAPF method derived from a Q-learning-based MARL framework. Initially, we introduce a personalized-enhanced communication mechanism based on dynamic graph topology, which ascertains the core aspects of ``who" and ``what" in interactive process through three-stage operations: selection, generation, and aggregation. Concurrently, we incorporate local crowd perception to enrich agents' heuristic observation, thereby strengthening the model's guidance for effective actions via the integration of static spatial constraints and dynamic occupancy changes. To resolve extreme deadlock issues, we propose a region-based deadlock-breaking strategy that leverages expert guidance to implement efficient coordination within confined areas. Experimental results demonstrate that PC2P achieves superior performance compared to state-of-the-art distributed MAPF methods in varied environments. Ablation studies further confirm the effectiveness of each module for overall performance.

多智能体路径规划强化学习协同决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。