用流匹配技术提升语音前端处理后的听感质量。
SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
- 基于条件流匹配,对前端处理后的语音进行后置优化。
- 在多种噪声和混响场景下显著提升听觉感知质量。
- 适合需要高质量语音输出的语音识别、语音合成等应用。
语音预处理技术如降噪、去混响和分离常用于各类下游语音任务中。然而,这些方法有时效果不佳,会残留噪声或引入新伪影,这类问题虽不被SI-SNR等指标捕捉,却能被人类听觉明显感知。为此,本文提出SpeechRefiner,一种利用条件流匹配(CFM)的后处理工具,用于提升语音的感知质量。我们在内部集成多个前端算法的处理流程中评估了SpeechRefiner,并与近期特定任务优化方法对比。实验表明,该方法在多样化的失真源下具有强泛化能力,显著改善语音听感。音频演示可访问 https://speechrefiner.github.io/SpeechRefiner/。
原文摘要 · Abstract (English)
Speech pre-processing techniques such as denoising, de-reverberation, and separation, are commonly employed as front-ends for various downstream speech processing tasks. However, these methods can sometimes be inadequate, resulting in residual noise or the introduction of new artifacts. Such deficiencies are typically not captured by metrics like SI-SNR but are noticeable to human listeners. To address this, we introduce SpeechRefiner, a post-processing tool that utilizes Conditional Flow Matching (CFM) to improve the perceptual quality of speech. In this study, we benchmark SpeechRefiner against recent task-specific refinement methods and evaluate its performance within our internal processing pipeline, which integrates multiple front-end algorithms. Experiments show that SpeechRefiner exhibits strong generalization across diverse impairment sources, significantly enhancing speech perceptual quality. Audio demos can be found at https://speechrefiner.github.io/SpeechRefiner/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。