用递归PPO让多架无人机在无GPS时精准定位目标
A Scalable Decentralized Reinforcement Learning Framework for UAV Target Localization Using Recurrent PPO
- 用带记忆的PPO算法让无人机自主学习定位
- 单机准确率93%,双机平均步数更少
- 适合复杂环境下的无人机群协同任务
无人机技术的快速发展推动了环境监测、灾害响应和农业调查等应用。提升多架去中心化无人机的协作能力,可显著增强这些应用的效率与协调性。本文研究了一种用于感知退化环境(如无GNSS/GPS信号区域)的目标定位的递归PPO模型。首先开发了单机目标识别方法,随后构建了去中心化的双机协同模型。该方法可利用无人机上的两种传感器:探测传感器和目标信号传感器。单机模型达到93%的准确率,双机模型准确率达86%,且平均定位步数更少。结果表明,该方法在无人机集群中具有潜力,可在复杂环境中高效、有效地实现辐射源目标的定位。
原文摘要 · Abstract (English)
The rapid advancements in unmanned aerial vehicles (UAVs) have unlocked numerous applications, including environmental monitoring, disaster response, and agricultural surveying. Enhancing the collective behavior of multiple decentralized UAVs can significantly improve these applications through more efficient and coordinated operations. In this study, we explore a Recurrent PPO model for target localization in perceptually degraded environments like places without GNSS/GPS signals. We first developed a single-drone approach for target identification, followed by a decentralized two-drone model. Our approach can utilize two types of sensors on the UAVs, a detection sensor and a target signal sensor. The single-drone model achieved an accuracy of 93%, while the two-drone model achieved an accuracy of 86%, with the latter requiring fewer average steps to locate the target. This demonstrates the potential of our method in UAV swarms, offering efficient and effective localization of radiant targets in complex environmental conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。