arXiv:2508.06840eess.ASeess.SP2025-08被引 34

用流匹配方法实现低计算量语音增强,推理仅需5次函数求值。

FlowSE: Flow Matching-based Speech Enhancement

  • 基于条件流匹配构建语音增强模型,替代传统扩散模型。
  • 在仅5次函数求值下性能接近60次求值的扩散模型。
  • 无需额外微调,适合实时语音处理场景。

扩散概率模型在语音增强任务中表现优异,但通常需要25至60次函数求值(NFE)进行推理,计算开销大。最近有方法通过微调修正反向过程,显著降低NFE。流匹配是一种训练连续归一化流的方法,可建模从已知分布到未知分布的概率路径,包括扩散过程描述的分布。本文提出一种基于条件流匹配的语音增强方法。该方法在NFE为5时,性能与NFE为60的扩散模型相当;在NFE为1至5时,性能与经微调的扩散模型相近,且无需额外微调过程。此外,我们还证明了由修改后的最优传输条件向量场导出的对应扩散模型,在无任何微调的情况下,于NFE=5时表现出类似性能。

原文摘要 · Abstract (English)

Diffusion probabilistic models have shown impressive performance for speech enhancement, but they typically require 25 to 60 function evaluations in the inference phase, resulting in heavy computational complexity. Recently, a fine-tuning method was proposed to correct the reverse process, which significantly lowered the number of function evaluations (NFE). Flow matching is a method to train continuous normalizing flows which model probability paths from known distributions to unknown distributions including those described by diffusion processes. In this paper, we propose a speech enhancement based on conditional flow matching. The proposed method achieved the performance comparable to those for the diffusion-based speech enhancement with the NFE of 60 when the NFE was 5, and showed similar performance with the diffusion model correcting the reverse process at the same NFE from 1 to 5 without additional fine tuning procedure. We also have shown that the corresponding diffusion model derived from the conditional probability path with a modified optimal transport conditional vector field demonstrated similar performances with the NFE of 5 without any fine-tuning procedure.

语音增强流匹配扩散模型低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。