基于子带编码与广义标签多伯努利滤波,提升语音基频估计鲁棒性。
Robust Pitch Estimation and Tracking for Speakers Based on Subband Encoding and the Generalized Labeled Multi-Bernoulli Filter
- 用听觉滤波器组分解信号,自适应确定子带数量和中心频率。
- 结合归一化自相关与新状态转移模型,显著降低误检率。
- 适合嘈杂或混响环境下的语音基频追踪,尤其对说话人识别有用。
本文提出一种新的基频估计算法与语音基频追踪方法。首先利用听觉滤波器组将声音信号分解为子带,假设人类语音具有时频稀疏性。不同于凭经验设定子带数,我们提出一种新颖的频率覆盖度量来自动推导子带数量及滤波器中心频率。子带信号通过受计算听觉场景分析(CASA)启发的方式编码,并计算归一化自相关以实现基频估计。为抑制虚假误差并跟踪说话人身份,引入时间连续性约束,采用广义标签多伯努利(GLMB)滤波器进行基频追踪,设计基于奥恩斯坦-乌伦贝克过程的新状态转移模型,以及基于测量驱动的新生目标模型以实现自适应新增目标。在多种加性噪声下的实验表明,所提方法在多数场景下优于多个当前最优基频估计算法;在混响房间的真实录音测试中也表现出良好的鲁棒性。
原文摘要 · Abstract (English)
This paper proposes a new pitch estimator and a novel pitch tracker for speakers. We first decompose the sound signal into subbands using an auditory filterbank, assuming time-frequency sparsity of human speech. Instead of directly selecting the number of subbands according to experience, we propose a novel frequency coverage metric to derive the number of subbands and the center frequencies of the filterbank. The subband signals are then encoded inspired by the computational auditory scene analysis (CASA) approach, and the normalized autocorrelations are calculated for pitch estimation. To suppress spurious errors and track the speaker identity, the temporal continuity constraint is exploited and a Generalized Labeled Multi-Bernoulli (GLMB) filter is adapted for pitch tracking, where we use a novel pitch state transition model based on the Ornstein-Uhlenbeck process, and the measurement driven birth model for adaptive new births of pitch targets. Experimental evaluations with various additive noises demonstrate that the proposed methods have achieved better accuracy compared with several state-of-the-art pitch estimation methods in most studied scenarios. Tests using real recordings in a reverberant room also show that the proposed method is robust against reverberation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。