用混合模型提升复杂声源定位精度,尤其在多声源环境下表现更优。
SHAMaNS: Sound Localization with Hybrid Alpha-Stable Spatial Measure and Neural Steerer
- 结合α稳定分布与神经波束成形器建模声波方向
- 在多声源场景下定位误差降低,优于现有方法
- 适合高噪声或非高斯干扰下的声源定位任务
本文提出一种声源定位(SSL)技术,将α-稳定分布模型与基于神经网络的波束成形器相结合。具体而言,采用物理信息引导的神经波束成形器(Neural Steerer)对固定麦克风阵列上的实测波束向量进行插值,从而更稳健地估计α-稳定空间度量,该度量代表目标信号最可能的到达方向(DOA)。由于α-稳定分布(α ∈ (0, 2))在非高斯情况下理论上定义唯一的空间度量,因此被用于建模神经波束成形器的残差重建误差,并在下游任务中加以利用。实验结果表明,该方法在多声源场景下显著优于当前最先进的技术。
原文摘要 · Abstract (English)
This paper describes a sound source localization (SSL) technique that combines an $α$-stable model for the observed signal with a neural network-based approach for modeling steering vectors. Specifically, a physics-informed neural network, referred to as Neural Steerer, is used to interpolate measured steering vectors (SVs) on a fixed microphone array. This allows for a more robust estimation of the so-called $α$-stable spatial measure, which represents the most plausible direction of arrival (DOA) of a target signal. As an $α$-stable model for the non-Gaussian case ($α$ $\in$ (0, 2)) theoretically defines a unique spatial measure, we choose to leverage it to account for residual reconstruction error of the Neural Steerer in the downstream tasks. The objective scores indicate that our proposed technique outperforms state-of-the-art methods in the case of multiple sound sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。