arXiv:2508.06310eess.AS2025-08被引 1

用混合模型提升无人机麦克风在强自噪声下的语音定位与增强效果

Egonoise Resilient Source Localization and Speech Enhancement for Drones Using a Hybrid Model and Learning-Based Approach

  • 结合阵列信号处理与深度神经网络,用六麦克风环形阵列实现语音增强
  • 在-30 dB 低信噪比下,性能优于四种基线方法
  • 适合无人机语音采集、搜救与军事应用中的声源定位场景

无人机在搜救与军事任务中日益重要,但其搭载的麦克风常受旋翼自噪声干扰,导致信噪比极低。本文提出一种混合方法,结合阵列信号处理(ASP)与深度神经网络(DNN),利用安装在四旋翼无人机上的六麦克风均匀环形阵列,实现目标说话人定位与语音增强。系统通过波束成形进行定位,并采用广义旁瓣对消器-深度滤波网络2(GSC-DF2)进行语音增强。使用DREGON数据集和实测数据验证,该方法在低至-30 dB的信噪比条件下,优于四种基线方法。

原文摘要 · Abstract (English)

Drones are becoming increasingly important in search and rescue missions, and even military operations. While the majority of drones are equipped with camera vision capabilities, the realm of drone audition remains underexplored due to the inherent challenge of mitigating the egonoise generated by the rotors. In this paper, we present a novel technique to address this extremely low signal-to-noise ratio (SNR) problem encountered by the microphone-embedded drones. The technique is implemented using a hybrid approach that combines Array Signal Processing (ASP) and Deep Neural Networks (DNN) to enhance the speech signals captured by a six-microphone uniform circular array mounted on a quadcopter. The system performs localization of the target speaker through beamsteering in conjunction with speech enhancement through a Generalized Sidelobe Canceller-DeepFilterNet 2 (GSC-DF2) system. To validate the system, the DREGON dataset and measured data are employed. Objective evaluations of the proposed hybrid approach demonstrated its superior performance over four baseline methods in the SNR condition as low as -30 dB.

语音增强无人机音频低信噪比声源定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。