首个三模态无人机跟踪数据集,解决复杂环境下多目标追踪难题
A Tri-Modal Dataset and a Baseline System for Tracking Unmanned Aerial Vehicles
- 融合可见光、红外与事件信号三模态数据,提升复杂场景追踪鲁棒性
- 提出自适应对齐与动态融合模块,在2.8万帧以上数据上实现领先性能
- 适合做无人机监控、智能安防及多模态感知研究的学者和工程师
随着低空无人机普及,视觉多目标追踪成为关键安全技术,尤其在光照不足、背景杂乱、运动快速等复杂条件下仍需高鲁棒性。单模态追踪常失效,而多模态追踪受限于缺乏专用公开数据集。为此,我们发布MM-UAV——首个大规模多模态无人机追踪基准,集成RGB、红外(IR)与事件信号三类感知模态,涵盖30多个挑战性场景,包含1,321个同步多模态序列和超过280万标注帧。配套提出专为无人机设计的多模态多目标追踪框架,包含两个关键技术:偏移引导的自适应对齐模块以解决传感器间空间错位,自适应动态融合模块平衡各模态互补信息;并引入事件增强关联机制,利用事件模态中的运动线索提升身份保持可靠性。大量实验表明,该框架持续优于现有先进方法。为推动研究,数据集与源代码将公开提供。
原文摘要 · Abstract (English)
With the proliferation of low altitude unmanned aerial vehicles (UAVs), visual multi-object tracking is becoming a critical security technology, demanding significant robustness even in complex environmental conditions. However, tracking UAVs using a single visual modality often fails in challenging scenarios, such as low illumination, cluttered backgrounds, and rapid motion. Although multi-modal multi-object UAV tracking is more resilient, the development of effective solutions has been hindered by the absence of dedicated public datasets. To bridge this gap, we release MM-UAV, the first large-scale benchmark for Multi-Modal UAV Tracking, integrating three key sensing modalities, e.g. RGB, infrared (IR), and event signals. The dataset spans over 30 challenging scenarios, with 1,321 synchronised multi-modal sequences, and more than 2.8 million annotated frames. Accompanying the dataset, we provide a novel multi-modal multi-UAV tracking framework, designed specifically for UAV tracking applications and serving as a baseline for future research. Our framework incorporates two key technical innovations, e.g. an offset-guided adaptive alignment module to resolve spatio mismatches across sensors, and an adaptive dynamic fusion module to balance complementary information conveyed by different modalities. Furthermore, to overcome the limitations of conventional appearance modelling in multi-object tracking, we introduce an event-enhanced association mechanism that leverages motion cues from the event modality for more reliable identity maintenance. Comprehensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art methods. To foster further research in multi-modal UAV tracking, both the dataset and source code will be made publicly available at https://xuefeng-zhu5.github.io/MM-UAV/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。