用非负矩阵分解提升混响语音清晰度,单通道下效果优于传统方法
Single Channel Blind Dereverberation of Speech Signals
- 基于非负矩阵分解的激活矩阵建模,从混响语音中重建干净语音频谱
- 在TIMIT数据集和Reverb2014数据上,新方法使PESQ提升0.15,对数谱失真降低1.8%
- 适合语音增强、语音识别前处理场景,尤其适用于单麦克风环境
录音语音的混响消除是语音处理中的核心问题。本文提出一种从混响语音频谱中估计干净语音频谱的方法,采用非负矩阵因子分解(NMFD)实现。进一步结合语音幅度谱的NMF表示,并引入卷积型NMF与帧堆叠模型以捕捉时序依赖性。提出一种新方法,将NMFD应用于混响语音幅度谱的激活矩阵进行去混响。基于TIMIT语料库的句子录音及Reverb 2014挑战中的实测房间冲激响应,通过客观指标PESQ和倒谱失真对多种技术进行了对比分析。尽管定性上验证了文献结论,但定量结果未能完全复现。所提新方法在量化指标上有所提升,但表现不够稳定。
原文摘要 · Abstract (English)
Dereverberation of recorded speech signals is one of the most pertinent problems in speech processing. In the present work, the objective is to understand and implement dereverberation techniques that aim at enhancing the magnitude spectrogram of reverberant speech signals to remove the reverberant effects introduced. An approach to estimate a clean speech spectrogram from the reverberant speech spectrogram is proposed. This is achieved through non-negative matrix factor deconvolution(NMFD). Further, this approach is extended using the NMF representation for speech magnitude spectrograms. To exploit temporal dependencies, a convolutive NMF-based representation and a frame-stacked model are incorporated into the NMFD framework for speech. A novel approach for dereverberation by applying NMFD to the activation matrix of the reverberated magnitude spectrogram is also proposed. Finally, a comparative analysis of the performance of the listed techniques, using sentence recordings from the TIMIT database and recorded room impulse responses from the Reverb 2014 challenge, is presented based on two key objective measures - PESQ and Cepstral Distortion.\\ Although we were qualitatively able to verify the claims made in literature regarding these techniques, exact results could not be matched. The novel approach, as it is suggested, provides improvement in quantitative metrics, but is not consistent
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。