arXiv:2508.04887eess.AS2025-08被引 1

提出闭式解方法,实现多声源逐个激活时的高效鲁棒波束成形

Closed-Form Successive Relative Transfer Function Vector Estimation based on Blind Oblique Projection Incorporating Noise Whitening

  • 用闭式解替代迭代优化,大幅降低计算开销
  • 采用正交附加向量,提升声源方向估计精度
  • 结合噪声白化技术,显著增强低信噪比下性能

相对传输函数(RTFs)在波束成形中至关重要,有助于有效抑制噪声与干扰。本文针对多个声源依次激活的场景,在混响和噪声环境下在线估计各源的RTF向量。虽然首个声源的RTF可直接估计,但后续声源在多源共存时段的估计面临挑战。现有盲斜投影(BOP)方法虽能估计新激活源的RTF,但存在计算复杂度高、依赖随机附加向量影响性能、需高信噪比等缺陷。为此,本文提出三项改进:首先推导出BOP代价函数的闭式解,显著降低计算复杂度;其次引入正交附加向量替代随机向量,提升估计精度;第三,融合协方差相减与噪声白化技术,增强低信噪比下的鲁棒性。为支持帧级源活动模式估计,还提出基于空间相干性的在线源数检测方法。仿真使用真实混响噪声录音,包含3个依次激活说话人,对比了有无先验源活动信息的情况。

原文摘要 · Abstract (English)

Relative transfer functions (RTFs) of sound sources play a crucial role in beamforming, enabling effective noise and interference suppression. This paper addresses the challenge of online estimating the RTF vectors of multiple sound sources in noisy and reverberant environments, for the specific scenario where sources activate successively. While the RTF vector of the first source can be estimated straightforwardly, the main challenge arises in estimating the RTF vectors of subsequent sources during segments where multiple sources are simultaneously active. The blind oblique projection (BOP) method has been proposed to estimate the RTF vector of a newly activating source by optimally blocking this source. However, this method faces several limitations: high computational complexity due to its reliance on iterative gradient descent optimization, the introduction of random additional vectors, which can negatively impact performance, and the assumption of high signal-to-noise ratio (SNR). To overcome these limitations, in this paper we propose three extensions to the BOP method. First, we derive a closed-form solution for optimizing the BOP cost function, significantly reducing computational complexity. Second, we introduce orthogonal additional vectors instead of random vectors, enhancing RTF vector estimation accuracy. Third, we incorporate noise handling techniques inspired by covariance subtraction and whitening, increasing robustness in low SNR conditions. To provide a frame-by-frame estimate of the source activity pattern, required by both the conventional BOP method and the proposed method, we propose a spatial-coherence-based online source counting method. Simulations are performed with real-world reverberant noisy recordings featuring 3 successively activating speakers, with and without a-priori knowledge of the source activity pattern.

波束成形声源定位信号处理低信噪比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。