arXiv:2604.01524eess.AS2026-04被引 14

在强混响下精准定位说话人,提升语音识别实际应用效果。

Reverberation-Robust Localization of Speakers Using Distinct Speech Onsets and Multi-channel Cross-Correlations

  • 通过听觉滤波器组分解信号,结合语音起始点检测与多通道互相关
  • 在混响时间达1秒的模拟和真实环境均实现稳定定位
  • 适合需要高鲁棒性的语音会议、智能音箱等场景

现有大量说话人定位方法,但在强混响环境下仍具挑战。本文提出两种基于麦克风阵列录音的定位算法。第一种方法利用听觉滤波器组将麦克风信号在时频域分解为子带,提出一种基于语音信号与冲击响应模型的新语音起始点检测方法,并构建各子带编码起始点的多通道互相关系数(MCCC),融合子带结果估计说话人方向。第二种方法扩展广义互相关-相位变换(GCC-PHAT)法,利用多麦克风冗余信息缓解混响影响。所提方法在仿真信号(混响时间 $T_{60}$ 达 1 秒)及真实混响房间($T_{60} \≈ 0.65$s)中测试,实验结果表明,相比部分先进方法,本方法可可靠定位静止与移动说话人。

原文摘要 · Abstract (English)

Many speaker localization methods can be found in the literature. However, speaker localization under strong reverberation still remains a major challenge in the real-world applications. This paper proposes two algorithms for localizing speakers using microphone array recordings of reverberated sounds. To separate concurrent speakers, the first algorithm decomposes microphone signals spectrotemporally into subbands via an auditory filterbank. To suppress reverberation, we propose a novel speech onset detection approach derived from the speech signal and impulse response models, and further propose to formulate the multi-channel cross-correlation coefficient (MCCC) of encoded speech onsets in each subband. The subband results are combined to estimate the directions-of-arrival (DOAs) of speakers. The second algorithm extends the generalized cross-correlation - phase transform (GCC-PHAT) method by using redundant information of multiple microphones to address the reverberation problem. The proposed methods have been evaluated under adverse conditions using not only simulated signals (reverberation time $T_{60}$ of up to $1$s) but also recordings in a real reverberant room ($T_{60} \approx 0.65$s). Comparing with some state-of-the-art localization methods, experimental results confirm that the proposed methods can reliably locate static and moving speakers, in presence of reverberation.

说话人定位混响抑制麦克风阵列语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。