分离语音与噪声的潜在表示能显著提升单通道语音增强效果
Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
- 用预训练VAE分别提取语音和噪声的潜在特征
- 清晰分离的潜在空间使语音增强性能优于标准VAE
- 适合研究语音增强中表征学习的学者参考
最近提出了一种基于变分自编码器(VAE)的单通道语音增强系统,采用贝叶斯排列训练方法,利用两个预训练的VAE获取语音和噪声的潜在表示。在此基础上,一个含噪VAE学习从含噪语音中生成语音与噪声的潜在表示以实现语音增强。修改预训练VAE的损失项会影响其语音与噪声的潜在表示。本文研究了不同潜在表示对语音增强性能的影响。在DNS3、WSJ0-QUT和VoiceBank-DEMAND数据集上的实验表明,语音与噪声表示在潜在空间中明显分离时,性能显著优于标准VAE所产生的重叠表示。
原文摘要 · Abstract (English)
Recently, a variational autoencoder (VAE)-based single-channel speech enhancement system using Bayesian permutation training has been proposed, which uses two pretrained VAEs to obtain latent representations for speech and noise. Based on these pretrained VAEs, a noisy VAE learns to generate speech and noise latent representations from noisy speech for speech enhancement. Modifying the pretrained VAE loss terms affects the pretrained speech and noise latent representations. In this paper, we investigate how these different representations affect speech enhancement performance. Experiments on the DNS3, WSJ0-QUT, and VoiceBank-DEMAND datasets show that a latent space where speech and noise representations are clearly separated significantly improves performance over standard VAEs, which produce overlapping speech and noise representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。