arXiv:2502.20326cs.ROcs.AI2025-02被引 11

用深度强化学习让多无人机在无卫星信号的室内自主搜救,实测表现优异。

Deep Reinforcement Learning based Autonomous Decision-Making for Cooperative UAVs: A Search and Rescue Real World Application

  • 用APF奖励函数优化策略,实现平滑安全的飞行轨迹。
  • 通过图注意力网络实时分配任务,计算开销极低且接近最优。
  • 融合深度相机与惯性数据抑制垂直漂移,厘米级高度稳定。

本文提出首个端到端框架,实现多架无人机在无全球导航卫星系统(GNSS)的室内环境中自主执行搜索与救援(SAR)任务,集成引导、导航与集中式任务分配。采用双延迟深度确定性策略梯度(Twin Delayed Deep Deterministic Policy Gradient)控制器,结合人工势场(APF)奖励函数,融合吸引与排斥势能,实现连续控制,加速收敛并生成更平滑、更安全的轨迹,优于仅基于距离的基线方法。任务协同分配通过深度图注意力网络(GAT)解决,在每个决策步骤上对无人机-任务图进行推理,实现近似最优分配,且对机载计算资源消耗极小。为抑制室内LiDAR-SLAM常见的Z轴漂移,采用轻量级互补滤波器融合深度相机高度与IMU垂直速度,无需外部信标即可达到厘米级高度稳定性。该系统部署于两架1米级四旋翼无人机,在专为北约Sapience自主协作无人机竞赛设计的复杂多层灾害模拟场景中完成飞行测试。相较于以往主要停留在仿真阶段的深度强化学习引导方法,本框架成功实现在复杂室内环境中的自主导航,并在2024年赛事中夺得第一名。结果表明,基于APF的DRL与GAT驱动的协作机制可有效转化为可靠的现实世界搜救应用。

原文摘要 · Abstract (English)

This paper presents the first end-to-end framework that combines guidance, navigation, and centralised task allocation for multiple UAVs performing autonomous search-and-rescue (SAR) in GNSS-denied indoor environments. A Twin Delayed Deep Deterministic Policy Gradient controller is trained with an Artificial Potential Field (APF) reward that blends attractive and repulsive potentials with continuous control, accelerating convergence and yielding smoother, safer trajectories than distance-only baselines. Collaborative mission assignment is solved by a deep Graph Attention Network that, at each decision step, reasons over the drone-task graph to produce near-optimal allocations with negligible on-board compute. To arrest the notorious Z-drift of indoor LiDAR-SLAM, we fuse depth-camera altimetry with IMU vertical velocity in a lightweight complementary filter, giving centimetre-level altitude stability without external beacons. The resulting system was deployed on two 1m-class quad-rotors and flight-tested in a cluttered, multi-level disaster mock-up designed for the NATO-Sapience Autonomous Cooperative Drone Competition. Compared with prior DRL guidance that remains largely in simulation, our framework demonstrates an ability to navigate complex indoor environments, securing first place in the 2024 event. These results demonstrate that APF-shaped DRL and GAT-driven cooperation can translate to reliable real-world SAR operations.

无人机强化学习搜救多机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。