用小阵列麦克风+卷积神经网络,实现1kHz以下高精度声源定位。
Adaptive high-precision sound source localization at low frequencies based on convolutional neural network

- 用小阵列压力分布输入,自定义标签与损失函数训练模型。
- 在低频段(<1kHz)定位误差显著低于传统波束成形、CLEAN-SC和DAMAS。
- 适应不同声源数和麦克风到网格距离,适合实际复杂场景应用。
声源定位(SSL)技术在故障诊断、语音分离和振动噪声抑制等领域至关重要。尽管波束成形算法广泛应用,但其在低频段的分辨率受限。近年来,基于深度学习的SSL方法通过使用大型麦克风阵列和特定训练数据提升了精度,但适用性较窄。为此,本文提出一种基于卷积神经网络的高精度低频声源定位方法,可在1kHz以下频率范围内,适应不同声源数量及麦克风阵列到扫描网格距离的变化。该方法以小型麦克风阵列上的压力分布为输入,采用定制化训练标签和损失函数进行模型训练。通过随机生成的测试数据集,在特定信噪比(SNR)条件下评估了模型的预测精度、适应性和鲁棒性,并与经典波束成形、CLEAN-SC和DAMAS算法进行对比。平面与空间声源分布结果均表明,所提神经网络模型显著提升了低频定位精度,展现出在SSL中的有效性与潜力。
原文摘要 · Abstract (English)
Sound source localization (SSL) technology plays a crucial role in various application areas such as fault diagnosis, speech separation, and vibration noise reduction. Although beamforming algorithms are widely used in SSL, their resolution at low frequencies is limited. In recent years, deep learning-based SSL methods have significantly improved their accuracy by employing large microphone arrays and training case specific neural networks, however, this could lead to narrow applicability. To address these issues, this paper proposes a convolutional neural network-based method for high-precision SSL, which is adaptive in the lower frequency range under 1kHz with varying numbers of sound sources and microphone array-to-scanning grid distances. It takes the pressure distribution on a relatively small microphone array as input to the neural network, and employs customized training labels and loss function to train the model. Prediction accuracy, adaptability and robustness of the trained model under certain signal-to-noise ratio (SNR) are evaluated using randomly generated test datasets, and compared with classical beamforming algorithms, CLEAN-SC and DAMAS. Results of both planar and spatial sound source distributions show that the proposed neural network model significantly improves low-frequency localization accuracy, demonstrating its effectiveness and potential in SSL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。