用仿生模型+强化学习,让无人船更灵活地拦截高机动敌船。
ARBoids: Adaptive Residual Reinforcement Learning With Boids Model for Cooperative Multi-USV Target Defense
- 结合鸟类群体行为模型与深度强化学习,实现多无人船协同防御。
- 在仿真中优于纯力学模型和普通强化学习策略,对不同机动性敌人有效。
- 适合研究智能舰艇对抗、多智能体协同控制的科研人员参考。
无人水面艇(USV)目标防御问题(TDP)旨在利用一个或多个防御型USV,在敌方USV突破指定目标区域前将其拦截。当攻击者具备远超防御者的机动能力时,该问题尤为棘手。为此,本文提出ARBoids:一种融合深度强化学习(DRL)与生物启发式、基于力的Boids模型的自适应残差强化学习框架。其中,Boids模型作为高效的基础协同策略,而DRL则学习残差策略以自适应优化防御者行动。该方法在高保真Gazebo仿真环境中验证,性能显著优于传统拦截策略,包括纯力控方法和基础DRL策略。所学策略对具有多样机动特性的攻击者表现出强适应性,证明其鲁棒性与泛化能力。代码将在论文接收后公开。
原文摘要 · Abstract (English)
The target defense problem (TDP) for unmanned surface vehicles (USVs) concerns intercepting an adversarial USV before it breaches a designated target region, using one or more defending USVs. A particularly challenging scenario arises when the attacker exhibits superior maneuverability compared to the defenders, significantly complicating effective interception. To tackle this challenge, this letter introduces ARBoids, a novel adaptive residual reinforcement learning framework that integrates deep reinforcement learning (DRL) with the biologically inspired, force-based Boids model. Within this framework, the Boids model serves as a computationally efficient baseline policy for multi-agent coordination, while DRL learns a residual policy to adaptively refine and optimize the defenders' actions. The proposed approach is validated in a high-fidelity Gazebo simulation environment, demonstrating superior performance over traditional interception strategies, including pure force-based approaches and vanilla DRL policies. Furthermore, the learned policy exhibits strong adaptability to attackers with diverse maneuverability profiles, highlighting its robustness and generalization capability. The code of ARBoids will be released upon acceptance of this letter.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。