让机器人团队在复杂环境里高效追捕,还能跨场景直接用。
Generalizable Collaborative Search-and-Capture in Cluttered Environments via Path-Guided MAPPO and Directional Frontier Allocation
- 用路径引导+前沿分配,让多机器人分工追捕更聪明。
- 在20x20新环境中零样本超越基线,抓人效率更高。
- 适合大规模机器人协同任务,部署简单不烧资源。
在杂乱环境中进行多智能体协同追捕面临奖励稀疏和视野受限的挑战。标准多智能体强化学习常因探索低效而难以扩展至大规模场景。本文提出PGF-MAPPO(路径引导前沿MAPPO),一种融合拓扑规划与反应式控制的分层框架。为解决局部最优和奖励稀疏问题,引入基于A*的势场实现密集奖励塑造;并提出方向性前沿分配策略,结合最远点采样(FPS)与几何角度抑制,强制空间分散并加速覆盖。架构采用参数共享的去中心化评判网络,保持O(1)模型复杂度,适合机器人集群应用。实验表明,PGF-MAPPO在对抗更快逃逸者时具备优异捕获效率;在10×10地图上训练的策略可实现零样本泛化至未见过的20×20环境,显著优于规则基和学习基基线方法。
原文摘要 · Abstract (English)
Collaborative pursuit-evasion in cluttered environments presents significant challenges due to sparse rewards and constrained Fields of View (FOV). Standard Multi-Agent Reinforcement Learning (MARL) often suffers from inefficient exploration and fails to scale to large scenarios. We propose PGF-MAPPO (Path-Guided Frontier MAPPO), a hierarchical framework bridging topological planning with reactive control. To resolve local minima and sparse rewards, we integrate an A*-based potential field for dense reward shaping. Furthermore, we introduce Directional Frontier Allocation, combining Farthest Point Sampling (FPS) with geometric angle suppression to enforce spatial dispersion and accelerate coverage. The architecture employs a parameter-shared decentralized critic, maintaining O(1) model complexity suitable for robotic swarms. Experiments demonstrate that PGF-MAPPO achieves superior capture efficiency against faster evaders. Policies trained on 10x10 maps exhibit robust zero-shot generalization to unseen 20x20 environments, significantly outperforming rule-based and learning-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。