用深度学习提升非平稳噪声下的语音存在概率估计精度
Learning-based A Posteriori Speech Presence Probability Estimation and Applications
- 融合全局与局部频段信息,增强网络对语音信号的感知能力
- 在非平稳噪声中实现更高信噪比下的语音增强效果
- 模型轻量,适合实时语音处理系统部署
后验语音存在概率(SPP)是噪声功率谱密度(PSD)估计的核心,对语音增强和语音识别系统至关重要。现有基于统计方法的SPP估计算法在非平稳噪声下性能受限,而深度学习方法常伴随高延迟。本文提出一种基于深度神经网络的改进SPP估计方法,在非平稳噪声条件下显著提升估计精度。通过编码器提取观测信号的全局信息,结合解码器与全连接层,利用残差连接中的混合全局-局部信息进行预测。为评估性能,采用基于当前帧SPP估计的次优最小均方误差(MMSE)方法直接估计噪声PSD,无需平滑处理,并以对数谱幅值估计器作为标准语音增强框架恢复纯净语音。实验表明,所提方法在保持低模型复杂度的同时,实现了更高的噪声PSD估计精度与语音增强性能。
原文摘要 · Abstract (English)
The a posteriori speech presence probability (SPP) is the fundamental component of noise power spectral density (PSD) estimation, which can contribute to speech enhancement and speech recognition systems. Most existing SPP estimators can estimate SPP accurately from the background noise. Nevertheless, numerous challenges persist, including the difficulty of accurately estimating SPP from non-stationary noise with statistics-based methods and the high latency associated with deep learning-based approaches. This paper presents an improved SPP estimation approach based on deep learning to achieve higher SPP estimation accuracy, especially in non-stationary noise conditions. To promote the information extraction performance of the DNN, the global information of the observed signal and the local information of the decoupled frequency bins from the observed signal are connected as hybrid global-local information. The global information is extracted by one encoder. Then, one decoder and two fully connected layers are used to estimate SPP from the information of residual connection. To evaluate the performance of our proposed SPP estimator, the noise PSD estimation and speech enhancement tasks are performed. In contrast to existing minimum mean-square error (MMSE)-based noise PSD estimation approaches, the noise PSD is estimated by the sub-optimal MMSE based on the current frame SPP estimate without smoothing. Directed by the noise PSD estimate, a standard speech enhancement framework, the log spectral amplitude estimator, is employed to extract clean speech from the observed signal. From the experimental results, we can confirm that our proposed SPP estimator can achieve high noise PSD estimation accuracy and speech enhancement performance while requiring low model complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。