用点引导自动生成无人机检测与分割标注,提升小目标识别能力。
UAVDB: Point-Guided Masks for UAV Detection and Segmentation
- 通过轨迹点生成边界框,免人工标注且定位精准
- 结合SAM2生成分割掩码,实现多任务标签丰富数据
- 覆盖从像素级到清晰可见的极端尺度变化,适合小目标研究
准确检测无人机对监控、安防和空域管理至关重要。然而现有数据集在规模、分辨率及捕捉极端尺度变化方面仍显不足。为此,我们提出UAVDB,一个基于点引导弱监督流程构建的无人机检测与分割基准数据集。引入轻量级标注方法Patch Intensity Convergence(PIC),将轨迹点转化为边界框,无需人工标注即可保持精确空间定位。在此基础上,利用SAM2生成分割掩码,丰富数据集的多任务标签。UAVDB包含固定摄像头多视角视频中的RGB帧,涵盖从清晰可见到接近单像素实例的各种条件下的无人机。定量结果显示,PIC结合SAM2在交并比(IoU)上优于现有标注技术。此外,我们在UAVDB上基准测试了基于YOLO的检测器,为后续研究建立基准。
原文摘要 · Abstract (English)
Accurate detection of Unmanned Aerial Vehicles (UAVs) is critical for surveillance, security, and airspace monitoring. However, existing datasets remain limited in scale, resolution, and the ability to capture objects across extreme size variations. To address these challenges, we present UAVDB, a benchmark dataset for UAV detection and segmentation, constructed via a point-guided weak supervision pipeline. We introduce Patch Intensity Convergence (PIC), a lightweight annotation method that converts trajectory points into bounding boxes, eliminating the need for manual labeling while preserving precise spatial localization. Building upon these annotations, we further generate segmentation masks using SAM2, enriching the dataset with multi-task labels. UAVDB consists of RGB frames from a fixed-camera multi-view video dataset, capturing UAVs across scales ranging from clearly visible objects to near single-pixel instances under diverse conditions. Quantitative results show that PIC combined with SAM2 outperforms existing annotation techniques in terms of IoU. Furthermore, we benchmark YOLO-based detectors on UAVDB, establishing baselines for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。