融合视觉与声音信息,提升无人机在复杂环境下的检测精度。
WAVE-DETR Multi-Modal Visible and Acoustic Real-Life Drone Detector
- 用门控机制融合视觉与声学特征,增强目标检测能力。
- 小无人机检测的mAP提升11.1%至15.3%,全尺寸无人机整体增益3.27%~5.84%。
- 适用于真实场景中低可见度或噪声干扰下的无人机监测任务。
本文提出一种结合可见光RGB与声学信号的多模态WAVE-DETR无人机检测器,用于提升真实场景下小型无人机的检测鲁棒性。该方法基于可变形DETR与Wav2Vec2架构,在现有Drone-vs-Bird数据集及新构建的包含超过7,500组同步图像与音频片段的ARDrone数据集上进行训练与测试。通过门控、线性层、MLP和交叉注意力四种融合方式,将Wav2Vec2提取的声学嵌入与可变形DETR的多尺度特征图进行融合,显著提升了不同尺寸无人机的检测性能。最优的门控融合策略在所有IoU阈值(0.5~0.9)下,使小无人机的mAP提升11.1%~15.3%,中大型无人机也获得3.27%~5.84%的整体增益。
原文摘要 · Abstract (English)
We introduce a multi-modal WAVE-DETR drone detector combining visible RGB and acoustic signals for robust real-life UAV object detection. Our approach fuses visual and acoustic features in a unified object detector model relying on the Deformable DETR and Wav2Vec2 architectures, achieving strong performance under challenging environmental conditions. Our work leverage the existing Drone-vs-Bird dataset and the newly generated ARDrone dataset containing more than 7,500 synchronized images and audio segments. We show how the acoustic information is used to improve the performance of the Deformable DETR object detector on the real ARDrone dataset. We developed, trained and tested four different fusion configurations based on a gated mechanism, linear layer, MLP and cross attention. The Wav2Vec2 acoustic embeddings are fused with the multi resolution feature mappings of the Deformable DETR and enhance the object detection performance over all drones dimensions. The best performer is the gated fusion approach, which improves the mAP of the Deformable DETR object detector on our in-distribution and out-of-distribution ARDrone datasets by 11.1% to 15.3% for small drones across all IoU thresholds between 0.5 and 0.9. The mAP scores for medium and large drones are also enhanced, with overall gains across all drone sizes ranging from 3.27% to 5.84%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。