arXiv:2509.23299cs.SDeess.AS2025-09被引 2

一歩生成式语音增强,速度更快、模型更小

MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow

  • 用平均速度场实现单步潜空间优化,加速推理
  • 在真实和模拟数据上达到顶尖听感质量,RTF更低
  • 基于自监督特征而非变分自编码器,适合实际部署

语音增强(SE)从噪声信号中恢复清晰语音,在通信和自动语音识别(ASR)中至关重要。虽然生成方法能提供优异的听感质量,但通常依赖多步采样(如扩散或流匹配)或大语言模型,限制了实时应用。为此,我们提出 MeanFlowSE,一种单步生成式语音增强框架。它采用 MeanFlow 预测平均速度场以实现单步潜空间修正,并以自监督学习(SSL)表示作为条件,而非变分自编码器(VAE)潜变量。该设计提升了推理速度,同时在训练中提供稳健的声学-语义引导。在 Interspeech 2020 DNS Challenge 盲测集和模拟测试集上,MeanFlowSE 达到顶尖听感质量,且在可比可懂度下显著降低实时因子(RTF)与模型规模,适合实际应用。代码将在发表后公开于 https://github.com/Hello3world/MeanFlowSE。

原文摘要 · Abstract (English)

Speech enhancement (SE) recovers clean speech from noisy signals and is vital for applications such as telecommunications and automatic speech recognition (ASR). While generative approaches achieve strong perceptual quality, they often rely on multi-step sampling (diffusion/flow-matching) or large language models, limiting real-time deployment. To mitigate these constraints, we present MeanFlowSE, a one-step generative SE framework. It adopts MeanFlow to predict an average-velocity field for one-step latent refinement and conditions the model on self-supervised learning (SSL) representations rather than VAE latents. This design accelerates inference and provides robust acoustic-semantic guidance during training. In the Interspeech 2020 DNS Challenge blind test set and simulated test set, MeanFlowSE attains state-of-the-art (SOTA) level perceptual quality and competitive intelligibility while significantly lowering both real-time factor (RTF) and model size compared with recent generative competitors, making it suitable for practical use. The code will be released upon publication at https://github.com/Hello3orld/MeanFlowSE.

语音增强生成模型单步生成实时部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。