通过局部信息聚合提升机器人集群动态任务分配的适应性与效率
A Local Information Aggregation based Multi-Agent Reinforcement Learning for Robot Swarm Dynamic Task Allocation
- 采用局部信息聚合模块,让机器人基于邻居数据协同决策
- 在复杂环境中实现快速适应,比六种传统算法更快收敛且更稳定
- 适合需要实时响应的多机器人协作场景,如搜救或巡检
本文研究动态环境下机器人集群的任务分配优化问题,强调构建鲁棒、灵活且可扩展的合作策略的重要性。提出一种基于分布式部分可观马尔可夫决策过程(Dec_POMDP)的新框架,核心为局部信息聚合多智能体深度确定性策略梯度(LIA_MADDPG)算法,采用集中训练、分布式执行(CTDE)机制。训练阶段引入局部信息聚合(LIA)模块,从邻近机器人收集关键信息以提升决策效率;执行阶段通过策略改进方法动态调整任务分配,应对环境变化与部分可观测性挑战。实验表明,LIA模块可无缝集成至多种基于CTDE的MARL方法中,显著提升性能。相比六种传统强化学习算法及一种启发式算法,LIA_MADDPG展现出更强的可扩展性、更快的环境适应速度,并保持高稳定性与收敛速度,验证了其在增强局部协作与自适应执行方面的卓越表现,具有显著提升机器人集群动态任务分配能力的潜力。
原文摘要 · Abstract (English)
In this paper, we explore how to optimize task allocation for robot swarms in dynamic environments, emphasizing the necessity of formulating robust, flexible, and scalable strategies for robot cooperation. We introduce a novel framework using a decentralized partially observable Markov decision process (Dec_POMDP), specifically designed for distributed robot swarm networks. At the core of our methodology is the Local Information Aggregation Multi-Agent Deep Deterministic Policy Gradient (LIA_MADDPG) algorithm, which merges centralized training with distributed execution (CTDE). During the centralized training phase, a local information aggregation (LIA) module is meticulously designed to gather critical data from neighboring robots, enhancing decision-making efficiency. In the distributed execution phase, a strategy improvement method is proposed to dynamically adjust task allocation based on changing and partially observable environmental conditions. Our empirical evaluations show that the LIA module can be seamlessly integrated into various CTDE-based MARL methods, significantly enhancing their performance. Additionally, by comparing LIA_MADDPG with six conventional reinforcement learning algorithms and a heuristic algorithm, we demonstrate its superior scalability, rapid adaptation to environmental changes, and ability to maintain both stability and convergence speed. These results underscore LIA_MADDPG's outstanding performance and its potential to significantly improve dynamic task allocation in robot swarms through enhanced local collaboration and adaptive strategy execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。