用吉布斯随机场建模多机器人编队,实现高成功率分布式控制。
Learning Efficient Flocking Control based on Gibbs Random Fields
- 基于吉布斯随机场设计奖励函数,从概率分布视角优化编队策略。
- 在复杂环境中实现近99%的成功率,优于现有主流方法。
- 适合需要高可靠性与可扩展性的多机器人协同场景。
编队控制对多机器人系统在各类应用中至关重要,但在拥挤环境中实现高效编队仍面临计算负担重、性能最优性差和运动安全性不足的挑战。本文提出一种基于吉布斯随机场(Gibbs Random Fields, GRFs)的多智能体强化学习(MARL)框架,将多机器人系统建模为服从联合概率分布的随机变量集合,为编队奖励设计提供新视角。通过基于GRF的信用分配方法,实现去中心化训练与执行,显著提升MARL在机器人数量增加时的可扩展性。引入动作注意力模块,隐式预测邻近机器人的运动意图,缓解MARL中的非平稳性问题。在仿真与实验中,该框架在复杂环境中实现了约99%的成功率,优于当前先进方法。消融实验验证了各模块的有效性。
原文摘要 · Abstract (English)
Flocking control is essential for multi-robot systems in diverse applications, yet achieving efficient flocking in congested environments poses challenges regarding computation burdens, performance optimality, and motion safety. This paper addresses these challenges through a multi-agent reinforcement learning (MARL) framework built on Gibbs Random Fields (GRFs). With GRFs, a multi-robot system is represented by a set of random variables conforming to a joint probability distribution, thus offering a fresh perspective on flocking reward design. A decentralized training and execution mechanism, which enhances the scalability of MARL concerning robot quantity, is realized using a GRF-based credit assignment method. An action attention module is introduced to implicitly anticipate the motion intentions of neighboring robots, consequently mitigating potential non-stationarity issues in MARL. The proposed framework enables learning an efficient distributed control policy for multi-robot systems in challenging environments with success rate around $99\%$, as demonstrated through thorough comparisons with state-of-the-art solutions in simulations and experiments. Ablation studies are also performed to validate the efficiency of different framework modules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。