用U-Net分割声波图实现无人机360°定位,精度更高且适配多种麦克风阵列。
Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD
- 将声源定位转为球面语义分割,直接输出声源分布区域。
- 在真实无人机数据上实现±2.5°角精度,跨环境泛化能力强。
- 方法不依赖具体麦克风布局,可快速迁移到新设备。
本文提出一种基于U-Net的360°声源定位方法,将问题建模为球面语义分割任务。不同于传统回归方向角(DoA)的方式,该模型对延迟叠加(DAS) beamformed 音频图(方位角与俯仰角)进行区域分割,识别出活跃声源位置。使用自研24麦克风阵列采集真实无人机(DJI Air 3)飞行数据,同步获取GPS轨迹、全景视频与飞行日志,构建带二值标签掩码的训练集。通过频率域表示的声能图输入改进型U-Net,结合Tversky损失缓解类别不平衡问题。网络输出经质心后处理生成稳健的DoA估计。实验表明,该方法在不同环境中均保持高精度(±2.5°),且无需重新训练即可适应新麦克风配置。进一步在DCASE 2019 TAU空间声事件数据集上验证,该框架适用于多类声事件定位与检测(SELD),展现强泛化能力。
原文摘要 · Abstract (English)
We introduce a U-net model for 360° acoustic source localization formulated as a spherical semantic segmentation task. Rather than regressing discrete direction-of-arrival (DoA) angles, our model segments beamformed audio maps (azimuth & elevation) into regions of active sound presence. Using delay-and-sum (DAS) beamforming on a custom 24-microphone array, we generate signals aligned with drone GPS telemetry to create binary supervision masks. A modified U-Net, trained on frequency-domain representations of these maps, learns to identify spatially distributed source regions while addressing class imbalance via the Tversky loss. Because the network operates on beamformed energy maps, the approach is inherently array-independent and can adapt to different microphone configurations and can be transferred to different microphone configurations with minimal adaptation. The segmentation outputs are post-processed by computing centroids over activated regions, enabling robust DoA estimates. Our dataset includes real-world open-field recordings of a DJI Air 3 drone, synchronized with 360° video and flight logs across multiple dates and locations. Experimental results show that U-net generalizes across environments, providing improved angular precision, offering a new paradigm for dense spatial audio understanding beyond traditional Sound Source Localization (SSL). We additionally validate the same beamforming-plus-segmentation formulation on the DCASE 2019 TAU Spatial Sound Events benchmark, showing that the approach generalizes beyond drone acoustics to multiclass Sound Event Localization and Detection (SELD) scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。