从噪声谱图生成清晰语音波形,提升语音质量。
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
- 分两步预测:先估噪声谱,再恢复干净谱
- 在VoiceBank+DEMAND数据集上表现领先
- 无需完整相位信息仍效果良好,适合降噪场景
本文提出一种新型神经去噪声码器,可从含噪梅尔谱图生成清晰语音波形。该模型由两个部分组成:谱预测器和增强模块。谱预测器首先从输入的含噪梅尔谱图中预测噪声幅度与相位谱,随后增强模块从中恢复出干净的幅度与相位谱。最终通过逆短时傅里叶变换(iSTFT)重建清晰语音波形。所有操作均在帧级谱域完成,分别采用APNet声码器和MP-SENet语音增强模型作为两部分的主干网络。实验表明,在VoiceBank+DEMAND数据集上,所提方法优于现有神经声码器,达到当前最优性能。尽管输入梅尔谱图缺乏完整相位信息且部分幅度信息缺失,该模型仍能实现与多个先进语音增强方法相当的效果。
原文摘要 · Abstract (English)
This paper proposes a novel neural denoising vocoder that can generate clean speech waveforms from noisy mel-spectrograms. The proposed neural denoising vocoder consists of two components, i.e., a spectrum predictor and a enhancement module. The spectrum predictor first predicts the noisy amplitude and phase spectra from the input noisy mel-spectrogram, and subsequently the enhancement module recovers the clean amplitude and phase spectrum from noisy ones. Finally, clean speech waveforms are reconstructed through inverse short-time Fourier transform (iSTFT). All operations are performed at the frame-level spectral domain, with the APNet vocoder and MP-SENet speech enhancement model used as the backbones for the two components, respectively. Experimental results demonstrate that our proposed neural denoising vocoder achieves state-of-the-art performance compared to existing neural vocoders on the VoiceBank+DEMAND dataset. Additionally, despite the lack of phase information and partial amplitude information in the input mel-spectrogram, the proposed neural denoising vocoder still achieves comparable performance with the serveral advanced speech enhancement methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。