用多智能体强化学习协同规划异构机器人团队任务与路径。
Collaborative Task and Path Planning for Heterogeneous Robotic Teams using Multi-Agent PPO
- 基于多智能体PPO算法实现机器人协作规划。
- 在行星探测场景中实现近优解且支持在线重规划。
- 适合需要实时协同的复杂任务场景研究者。
高效的外星探测需依赖具备多样化能力的机器人,从科学测量工具到先进运动系统。机器人团队通过将任务分配给多个专业化子系统,可提升任务完成效率和科学价值提取。核心挑战在于如何高效协调团队,以最大化资源利用与科学产出。传统规划算法随问题规模增长而急剧恶化,导致规划周期长、推理成本高,因机器人-目标分配组合与轨迹空间呈组合爆炸。基于学习的方法将扩展性问题从运行时转移至训练阶段,为实现实时规划奠定基础。本文提出一种基于多智能体近端策略优化(MAPPO)的协同规划策略,用于解决异构机器人团队的复杂目标分配与调度问题。我们通过穷举搜索获得单目标最优解作为基准,并在行星探测场景中评估该方法的在线重规划能力。
原文摘要 · Abstract (English)
Efficient robotic extraterrestrial exploration requires robots with diverse capabilities, ranging from scientific measurement tools to advanced locomotion. A robotic team enables the distribution of tasks over multiple specialized subsystems, each providing specific expertise to complete the mission. The central challenge lies in efficiently coordinating the team to maximize utilization and the extraction of scientific value. Classical planning algorithms scale poorly with problem size, leading to long planning cycles and high inference costs due to the combinatorial growth of possible robot-target allocations and possible trajectories. Learning-based methods are a viable alternative that move the scaling concern from runtime to training time, setting a critical step towards achieving real-time planning. In this work, we present a collaborative planning strategy based on Multi-Agent Proximal Policy Optimization (MAPPO) to coordinate a team of heterogeneous robots to solve a complex target allocation and scheduling problem. We benchmark our approach against single-objective optimal solutions obtained through exhaustive search and evaluate its ability to perform online replanning in the context of a planetary exploration scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。