arXiv:2509.21522cs.SDcs.AI2025-09被引 1

单阶段训练实现语音增强的高效推理,一步完成音质媲美多步扩散模型。

Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training

  • 单阶段训练生成不变于步骤的流模型,支持任意步数推理。
  • 单步推断在消费级显卡上实现实时因子0.013,质量接近60步扩散模型。
  • 揭示随机性在训练与推理中的作用,平衡音质与延迟。

基于扩散的生成模型在语音增强(SE)中实现了顶尖的感知质量,但其迭代特性需大量神经函数评估(NFE),难以满足实时应用需求。相比之下,流匹配通过学习直接向量场,利用确定性常微分方程(ODE)求解器仅需少数几步即可实现高质量合成。本文提出语音增强的快捷流匹配(SFMSE),一种新颖方法,通过单阶段训练构建一个步骤无关的模型。在训练中对目标时间步进行条件化,使模型无需架构修改或微调即可实现单步、少步或多步去噪。实验表明,单步推理在消费级GPU上实现0.013的实时因子(RTF),同时感知质量可媲美需60次NFE的强扩散基线。本工作还对训练与推理中随机性的作用进行了实证分析,弥合了高质量生成式语音增强与低延迟约束之间的差距。

原文摘要 · Abstract (English)

Diffusion-based generative models have achieved state-of-the-art performance for perceptual quality in speech enhancement (SE). However, their iterative nature requires numerous Neural Function Evaluations (NFEs), posing a challenge for real-time applications. On the contrary, flow matching offers a more efficient alternative by learning a direct vector field, enabling high-quality synthesis in just a few steps using deterministic ordinary differential equation~(ODE) solvers. We thus introduce Shortcut Flow Matching for Speech Enhancement (SFMSE), a novel approach that trains a single, step-invariant model. By conditioning the velocity field on the target time step during a one-stage training process, SFMSE can perform single, few, or multi-step denoising without any architectural changes or fine-tuning. Our results demonstrate that a single-step SFMSE inference achieves a real-time factor (RTF) of 0.013 on a consumer GPU while delivering perceptual quality comparable to a strong diffusion baseline requiring 60 NFEs. This work also provides an empirical analysis of the role of stochasticity in training and inference, bridging the gap between high-quality generative SE and low-latency constraints.

语音增强流匹配低延迟扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。