用强化学习动态调整防御机器人队形,有效驱赶高机动攻击群
Task-Oriented Formation Decision via Reinforcement Learning: Herding an Attacking Swarm

- 用低维参数表示队形,实现形状自适应优化
- 仿真训练的策略在真实场景中成功驱赶3攻7防团队
- 适用于数十个机器人的大规模协同任务
多机器人系统可通过编排特定队形完成单个机器人难以实现的任务。与现有研究中关注几何形状形成的差异在于,本文聚焦于任务导向的队形决策问题,以驱赶任务为核心。该任务因攻击方机动性强且策略未知而极具挑战性。为此,我们提出两项创新:首先,采用低维参数向量编码队形形状,将队形决策转化为参数优化问题,突破预设形状的灵活性限制;通过优化参数,防御方能以动态适配的队形弥补机动性劣势。其次,设计基于强化学习的策略来调控队形参数,在覆盖多种攻击策略的离线仿真中训练,使策略在在线部署时有效应对敌方不可预测行为。对比三种基线方法的仿真表明,本方法可成功完成复杂驱赶任务;扩展性仿真验证其在数十个机器人场景下的适用性。此外,我们在包含3名攻击者和7名防御者的物理机器人平台上验证了方法的实际可行性。
原文摘要 · Abstract (English)
Multi-robot systems can accomplish tasks that are difficult for a single robot by organizing into task-specific formations. Different from existing studies on multi-robot shape formation, we here study the task-oriented formation decision problem, with a focus on the herding task. This task is challenging due to the attackers' superior maneuverability and their unknown strategies. To address these challenges, we propose the following novel results. First, we encode the formation shape using a low-dimensional parameter vector. This parametric representation reformulates the formation decision as a parameter optimization problem, thereby resolving the limited flexibility of predefined shapes. By optimizing these formation parameters, the defenders' maneuverability disadvantage is mitigated through a formation shape that continuously adapts to task requirements. Second, we develop a reinforcement learning-based policy to regulate the formation parameters. Trained offline in simulations covering diverse attacking strategies, the learned policy can effectively handle adversarial unpredictability during online deployment. Comparative simulations against three baselines demonstrate that our method can successfully accomplish challenging herding tasks. Additional scalability simulations further verify its applicability to simulated scenarios involving dozens of robots. We also validate the practical feasibility of our method on a physical robotic platform with 3 attackers and 7 defenders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。