arXiv:2607.05939cs.RO2026-07

用强化学习让无人机团队用网捕获敏捷目标,效果优于传统方法。

Intercepting an Agile Target with Net-Carrying Drones using Competitive Multi-Agent Reinforcement Learning

论文配图:Intercepting an Agile Target with Net-Carrying Drones using Competitive Multi-Agent Reinforcement Learning
图 1 · 摘自论文原文
  • 采用竞争性多智能体强化学习,结合优先虚构自演法优化策略。
  • 捕获率更高,耗时更短,且碰撞率低,显著优于启发式基准。
  • 能自动生成协作战术,适合复杂空战与无人机集群研究者。

本文提出一种利用携带捕网的敏捷无人机团队拦截机动目标的解决方案。将问题建模为竞争性多智能体强化学习(MARL)任务,采用带优先虚构自演法(PFSP)的多智能体近端策略优化(MAPPO)训练追击者与逃逸者。在高保真模拟器中使用低层控制指令(集体推力与机体速率,CTBR)进行训练,实现双方高机动飞行。对比基于启发式的基线方法,所提方案在捕获率、捕获时间及碰撞率上均表现更优。消融实验表明,PFSP提升了策略鲁棒性,使其能适应不同对手策略;低层控制指令对学习有效策略至关重要。最后的定性分析显示,追击者间涌现出协同作战行为。

原文摘要 · Abstract (English)

This article presents a solution to intercept an agile drone by a team of agile drone carrying catching nets. We formulate the problem as a competitive Multi-Agent Reinforcement Learning (MARL) task. To address the problem of nonstationarity and catastrophic forgetting of agents overfitting to the current opponent strategy, we train the pursuers and the evader using Multi-Agent Proximal Policy Optimization (MAPPO) with Prioritized Fictitious Self Play (PFSP). We train the agents in a high-fidelity simulator using low-level control commands, collective thrust and body rates (CTBR), to achieve agile flights for both the pursuers and the evader. We compare the performance of the trained policies in terms of catch rate, time to catch and crash rates, against heuristic baselines and show that our solution outperforms them. Ablation studies show that PFSP lead to more robust policies that can adapt to different opponent strategies, and that a low-level control commands are crucial for learning performing strategies in the pursuit-evasion task. Finally, a qualitative analysis of the learned behaviours highlights the emergence of cooperative tactics among the pursuers.

多智能体强化学习无人机拦截协同策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。