用分步方法修复噪声混响中的语音,提升听感与下游任务表现。
Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules
- 先降噪再用生成式编解码器去混响并恢复语音细节。
- 在恶劣环境下使语音质量显著提升,尤其在低信噪比场景下表现优异。
- 适合语音识别、语音合成等需要高质量输入的下游应用。
在强噪声和混响环境下,传统语音增强(SE)方法常导致语音过度抑制,产生听感瑕疵并影响下游任务性能。为此,我们提出一种新型方法Restorative SE(RestSE),结合轻量级SE模块与生成式编解码器模块,分步提升语音质量:首先由SE模块降噪,再由编解码器模块进行去混响与语音恢复。系统研究了编解码器中多种量化技术对性能的影响,并引入加权损失函数与特征融合策略,将SE输出与原始混合信号融合,尤其在SE输出严重失真的段落中有效改善表现。实验结果表明,该方法在恶劣环境下的语音增强效果显著优于现有方法。音频演示可访问:https://sophie091524.github.io/RestorativeSE/。
原文摘要 · Abstract (English)
In challenging environments with significant noise and reverberation, traditional speech enhancement (SE) methods often lead to over-suppressed speech, creating artifacts during listening and harming downstream tasks performance. To overcome these limitations, we propose a novel approach called Restorative SE (RestSE), which combines a lightweight SE module with a generative codec module to progressively enhance and restore speech quality. The SE module initially reduces noise, while the codec module subsequently performs dereverberation and restores speech using generative capabilities. We systematically explore various quantization techniques within the codec module to optimize performance. Additionally, we introduce a weighted loss function and feature fusion that merges the SE output with the original mixture, particularly at segments where the SE output is heavily distorted. Experimental results demonstrate the effectiveness of our proposed method in enhancing speech quality under adverse conditions. Audio demos are available at: https://sophie091524.github.io/RestorativeSE/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。