arXiv:2509.21214eess.AS2025-09被引 2

用均值流实现单次计算的高效语音增强,效果远超传统生成模型。

MeanSE: Efficient Generative Speech Enhancement with Mean Flows

  • 通过建模平均速度场,实现单次函数评估(1-NFE)的语音增强
  • 在单次计算下性能显著优于基线流匹配模型,提升明显
  • 特别适合对实时性要求高、需跨域泛化的语音处理场景

语音增强旨在改善降质语音的质量,生成式模型如流匹配因其出色的感知质量受到关注。然而,基于流的模型需要多次函数评估(NFE)才能达到稳定且满意的效果,导致计算开销大,且1-NFE性能较差。本文提出MeanSE,一种基于均值流的高效生成式语音增强模型,通过建模平均速度场,实现高质量的单次函数评估(1-NFE)增强。实验表明,所提方法在仅使用1次NFE时,显著优于流匹配基线,展现出极强的跨域泛化能力。

原文摘要 · Abstract (English)

Speech enhancement (SE) improves degraded speech's quality, with generative models like flow matching gaining attention for their outstanding perceptual quality. However, the flow-based model requires multiple numbers of function evaluations (NFEs) to achieve stable and satisfactory performance, leading to high computational load and poor 1-NFE performance. In this paper, we propose MeanSE, an efficient generative speech enhancement model using mean flows, which models the average velocity field to achieve high-quality 1-NFE enhancement. Experimental results demonstrate that our proposed MeanSE significantly outperforms the flow matching baseline with a single NFE, exhibiting extremely better out-of-domain generalization capabilities.

语音增强生成模型流匹配高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。