arXiv:2505.19476eess.ASeess.SP2025-05中稿 · InterSpeech 2025被引 25

用流匹配实现高效高质语音增强,推理快且效果优。

FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching

  • 基于流匹配,在单次前向传播中完成噪声到清晰语音的连续变换。
  • 在无文本时仍保持高质量重建,有文本时性能进一步提升。
  • 相比扩散模型更快,比语言模型更保真,适合实时语音增强场景。

生成模型在音频任务中表现优异,包括语言模型、扩散模型和流匹配。然而,现有的语音增强(SE)生成方法存在显著挑战:基于语言模型的方法因量化损失导致说话人相似性和可懂度下降;而扩散模型训练复杂且推理延迟高。为此,我们提出FlowSE,一种基于流匹配的语音增强模型。流匹配在单次前向传播中学习噪声与清晰语音分布间的连续变换,显著降低推理延迟并保持高质量重建。具体而言,FlowSE在含噪梅尔频谱图及可选字符序列上训练,以真实梅尔频谱图为监督优化条件流匹配损失,隐式学习语音的时频结构与文本-语音对齐关系。推理时,FlowSE可不依赖文本信息运行,也可结合字幕进一步提升性能。大量实验表明,FlowSE显著优于现有先进生成方法,为生成式语音增强树立新范式,展示了流匹配在该领域的潜力。代码、预训练检查点及音频样例已公开。

原文摘要 · Abstract (English)

Generative models have excelled in audio tasks using approaches such as language models, diffusion, and flow matching. However, existing generative approaches for speech enhancement (SE) face notable challenges: language model-based methods suffer from quantization loss, leading to compromised speaker similarity and intelligibility, while diffusion models require complex training and high inference latency. To address these challenges, we propose FlowSE, a flow-matching-based model for SE. Flow matching learns a continuous transformation between noisy and clean speech distributions in a single pass, significantly reducing inference latency while maintaining high-quality reconstruction. Specifically, FlowSE trains on noisy mel spectrograms and optional character sequences, optimizing a conditional flow matching loss with ground-truth mel spectrograms as supervision. It implicitly learns speech's temporal-spectral structure and text-speech alignment. During inference, FlowSE can operate with or without textual information, achieving impressive results in both scenarios, with further improvements when transcripts are available. Extensive experiments demonstrate that FlowSE significantly outperforms state-of-the-art generative methods, establishing a new paradigm for generative-based SE and demonstrating the potential of flow matching to advance the field. Our code, pre-trained checkpoints, and audio samples are available.

语音增强流匹配高效推理生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。