arXiv:2506.18264cs.RO2025-06中稿 · presentation in pr…

用强化学习让无人机高效跟踪空中目标,兼顾精度与实时性。

Learning Approach to Efficient Vision-based Active Tracking of a Flying Target by an Unmanned Aerial Vehicle

  • 融合深度学习与核相关滤波,提升检测效率与精度。
  • 强化学习训练神经控制器,跟踪时长更久且距离更近。
  • 适合需要自主追踪的无人机系统研发者参考。

自主追踪空中飞行物在搜救、反无人机等场景中有重要应用。地面追踪需部署基础设施,受距离和环境限制。通过另一架飞行无人机(如追击型UAV)进行视觉主动追踪可弥补此短板,并支持空中协同。该任务涉及两大耦合问题:1)高效准确的目标检测与状态估计;2)决策飞行姿态以确保目标持续在视场内并利于后续检测。针对第一问题,本文提出将标准深度学习架构与核相关滤波(KCF)结合,实现计算高效且不牺牲精度的检测,优于单一学习或滤波方法。感知框架在实验室规模系统中验证。针对第二问题,为突破传统控制器对线性假设和背景变化的依赖,提出使用强化学习训练神经控制器,以快速生成速度机动指令。设计了新状态空间、动作空间与奖励函数,并在AirSim中完成仿真训练。测试表明,该模型在复杂目标机动下表现优于基线PID控制,在跟踪时长和平均保持距离上均有提升。

原文摘要 · Abstract (English)

Autonomous tracking of flying aerial objects has important civilian and defense applications, ranging from search and rescue to counter-unmanned aerial systems (counter-UAS). Ground based tracking requires setting up infrastructure, could be range limited, and may not be feasible in remote areas, crowded cities or in dense vegetation areas. Vision based active tracking of aerial objects from another airborne vehicle, e.g., a chaser unmanned aerial vehicle (UAV), promises to fill this important gap, along with serving aerial coordination use cases. Vision-based active tracking by a UAV entails solving two coupled problems: 1) compute-efficient and accurate (target) object detection and target state estimation; and 2) maneuver decisions to ensure that the target remains in the field of view in the future time-steps and favorably positioned for continued detection. As a solution to the first problem, this paper presents a novel integration of standard deep learning based architectures with Kernelized Correlation Filter (KCF) to achieve compute-efficient object detection without compromising accuracy, unlike standalone learning or filtering approaches. The proposed perception framework is validated using a lab-scale setup. For the second problem, to obviate the linearity assumptions and background variations limiting effectiveness of the traditional controllers, we present the use of reinforcement learning to train a neuro-controller for fast computation of velocity maneuvers. New state space, action space and reward formulations are developed for this purpose, and training is performed in simulation using AirSim. The trained model is also tested in AirSim with respect to complex target maneuvers, and is found to outperform a baseline PID control in terms of tracking up-time and average distance maintained (from the target) during tracking.

无人机追踪强化学习视觉跟踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。