提出分阶段表征学习框架,提升无人机实时追踪鲁棒性
Progressive Representation Learning for Real-Time UAV Tracking
- 分两阶段学习:先粗后细,结合外观与语义信息增强表征
- 在三个主流无人机追踪数据集上表现优异,实时性达42.6帧/秒
- 适合边缘设备部署,适用于复杂动态环境下的无人机应用
视觉目标追踪显著推动了无人飞行器(UAV)的自主应用。然而,在复杂动态环境中,面对尺度变化和遮挡问题,学习鲁棒的目标表征尤为困难,这些挑战会严重改变目标的原始信息。为此,本文提出一种新的分阶段表征学习框架——PRL-Track。该框架分为粗粒度表征学习和细粒度表征学习两个阶段。在粗粒度阶段,设计两种依赖外观和语义信息的调节器,以减轻外观干扰并捕捉语义信息;在细粒度阶段,引入新型层次化建模生成器,融合粗粒度目标表征。大量实验表明,所提PRL-Track在三个权威的无人机追踪基准上表现卓越。实际平台测试显示,其在配备边缘智能摄像头的典型UAV平台上实现42.6帧/秒的追踪速度。代码、模型及演示视频已开源。
原文摘要 · Abstract (English)
Visual object tracking has significantly promoted autonomous applications for unmanned aerial vehicles (UAVs). However, learning robust object representations for UAV tracking is especially challenging in complex dynamic environments, when confronted with aspect ratio change and occlusion. These challenges severely alter the original information of the object. To handle the above issues, this work proposes a novel progressive representation learning framework for UAV tracking, i.e., PRL-Track. Specifically, PRL-Track is divided into coarse representation learning and fine representation learning. For coarse representation learning, two innovative regulators, which rely on appearance and semantic information, are designed to mitigate appearance interference and capture semantic information. Furthermore, for fine representation learning, a new hierarchical modeling generator is developed to intertwine coarse object representations. Exhaustive experiments demonstrate that the proposed PRL-Track delivers exceptional performance on three authoritative UAV tracking benchmarks. Real-world tests indicate that the proposed PRL-Track realizes superior tracking performance with 42.6 frames per second on the typical UAV platform equipped with an edge smart camera. The code, model, and demo videos are available at \url{https://github.com/vision4robotics/PRL-Track}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。