arXiv:2509.15702eess.AScs.SD2025-09中稿 · publication in IEE…被引 1

改进声源定位方法,能适应复杂真实声学环境。

A Steered Response Power Method for Sound Source Localization With Generic Acoustic Models

  • 提出通用声学模型下的广义波束形成准则,支持任意麦克风布局。
  • 在噪声环境下定位误差降低超60%,显著优于传统SRP方法。
  • 适合语音识别、智能音箱等需高精度定位的场景。

传统的定向响应功率(SRP)方法常依赖理想声学假设,如全向远场声源、自由场传播和不相关噪声。然而现实中这些假设常被违反。本文提出一种通用化SRP方法,可适配任意麦克风配置与任意声学模型,包括分布式麦克风的幅值差异、声源/接收器指向性及声学遮蔽效应,并支持实测传递函数作为模型输入。研究表明,传统延迟求和波束形成不再适用于通用模型。为此,本文推导出考虑通用声学模型与空间相关噪声的最优SRP波束形成器,并设计合适的频率加权策略。新方法可联合利用麦克风信号间的幅值差与时间差进行定位。三种不同麦克风阵列在多种噪声条件下的仿真结果表明,该方法相比传统SRP显著降低平均定位误差,在噪声条件下误差减少超过60%。

原文摘要 · Abstract (English)

The steered response power (SRP) method is one of the most popular approaches for acoustic source localization with microphone arrays. It is often based on simplifying acoustic assumptions, such as an omnidirectional sound source in the far field of the microphone array(s), free field propagation, and spatially uncorrelated noise. In reality, however, there are many acoustic scenarios where such assumptions are violated. This paper proposes a generalization of the conventional SRP method that allows to apply generic acoustic models for localization with arbitrary microphone constellations. These models may consider, for instance, level differences in distributed microphones, the directivity of sources and receivers, or acoustic shadowing effects. Moreover, also measured acoustic transfer functions may be applied as acoustic model. We show that the delay-and-sum beamforming of the conventional SRP is not optimal for localization with generic acoustic models. To this end, we propose a generalized SRP beamforming criterion that considers generic acoustic models and spatially correlated noise, and derive an optimal SRP beamformer. Furthermore, we propose and analyze appropriate frequency weightings. Unlike the conventional SRP, the proposed method can jointly exploit observed level and time differences between the microphone signals to infer the source location. Realistic simulations of three different microphone setups with speech under various noise conditions indicate that the proposed method can significantly reduce the mean localization error compared to the conventional SRP and, in particular, a reduction of more than 60% can be archived in noisy conditions.

声源定位麦克风阵列波束形成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。