提出FocalSE,在编码器嵌入空间中实现降噪与噪声分离,提升低比特率下的语音重建质量。
Noisy Environment Adaptation of Neural Speech Codec via Focal Mask and Noise Feature Separation

- 在神经语音编码器的连续嵌入空间中进行特征去噪与噪声分离。
- 采用焦点调制压缩解压,有效捕捉全局上下文与局部互信息。
- 结合1D ResNet识别噪声类别,增强不同噪声场景下的分离效果。
神经语音编码器因其在低比特率下实现高质量语音重建而受到广泛关注。然而,真实环境中的噪声严重降低其性能,阻碍高质量纯净语音的重建。为解决此问题,我们提出FocalSE,一种新型语音增强方法,在神经语音编码器的连续嵌入空间中同时实现特征去噪、噪声特征分离与噪声识别。具体地,我们设计基于焦点调制的压缩与解压机制,以捕捉全局上下文和局部互信息,并生成焦点掩码以恢复纯净特征嵌入。随后,将噪声嵌入从含噪嵌入中分离,提升去噪性能。最后,使用ResNet1D-18识别噪声类别,以增强分离效果。在两个标准数据集LibriTTS和ESC50上的大量实验表明,本方法在低比特率与低信噪比条件下均优于现有先进方法。
原文摘要 · Abstract (English)
Neural speech codec has attracted extensive attention for high-quality reconstruction at low-bitrate. However, real-world noise severely degrades its performance and hinders high-quality clean speech reconstruction. To tackle this problem, we propose FocalSE, a novel speech enhancement method that performs feature denoising, noise feature separation and noise recognition in the continuous embedding space of neural speech codecs. Specifically, we develop focal modulation-based compression and decompression to capture global context and local mutual information, and generate focal masks to recover clean feature embeddings. We then separate noise embeddings from noisy embeddings to improve denoising performance. Finally, we use ResNet1D-18 to recognize noise categories for better separation effectiveness. Extensive experiments on two standard datasets, LibriTTS and ESC50, demonstrate that our method outperforms state-of-the-art approaches under low-bitrate and low-SNR conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。