用语言模型提升语音增强的语义一致性,让合成语音更自然可信。
SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement
- 分两阶段设计,用语言模型建模语义先验指导生成
- 在多种噪声下性能领先,尤其在复杂失真时表现突出
- 适合需要高保真语音还原的语音识别与通信场景
生成式通用语音增强方法旨在利用生成模型改善各类失真条件下的语音质量。然而,现有方法常因生成结果语义不一致而受限。为此,我们提出SenSE,一种基于语言模型建模语义先验的两阶段生成式通用语音增强框架,通过流匹配机制引导生成语义忠实的语音,显著提升上下文保真度。此外,引入双路径掩码条件训练策略,使流匹配增强过程可灵活融合来自降噪语音、语义标记和参考语音的多源条件信号,增强模型灵活性与适应性。实验表明,SenSE在生成式语音增强模型中达到最先进水平,尤其在挑战性失真条件下展现出极高的性能上限。代码与演示见https://github.com/ASLP-lab/SenSE。
原文摘要 · Abstract (English)
Generative Universal Speech Enhancement (USE) methods aim to leverage generative models to improve speech quality under various types of distortions. However, existing generative speech enhancement methods often suffer from semantic inconsistency in the generated outputs. Therefore, we propose SenSE, a novel two-stage generative universal speech enhancement framework, by modeling semantic priors with a language model, the flow matching-based speech enhancement process is guided to generate semantically faithful speech, thereby effectively improving context fidelity. In addition, we introduce a dual-path masked conditioning training strategy that enables flow matching-based enhancement to flexibly integrate multi-source conditioning signals from degraded speech, semantic tokens, and reference speech, thereby improving model flexibility and adaptability. Experimental results demonstrate that SenSE achieves state-of-the-art performance among generative speech enhancement models and exhibits a high performance ceiling, particularly under challenging distortion conditions. Codes and demos are available at https://github.com/ASLP-lab/SenSE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。