arXiv:2510.16997eess.AS2025-10被引 3

用流匹配实现20毫秒低延迟语音修复,支持实时通信。

Towards Real-Time Generative Speech Restoration with Flow-Matching

  • 采用因果架构无时间下采样,实现20毫秒总延迟。
  • 仅5次函数求值即达高还原质量,性能接近20次采样。
  • 适合实时语音通话,对小模型更优少步数采样。

生成模型在语音增强与修复任务中表现优异,但多数方法为离线处理,延迟高,不适用于流式应用。本文研究基于流匹配(Flow-Matching, FM)的低延迟、实时生成语音修复系统可行性。所提方法可应对真实场景中的降噪、去混响及生成式修复任务。采用无时间下采样的因果架构,总延迟仅20毫秒,适合实时通信。我们探索多种结构变体与采样策略,以保障训练有效性和推理效率。值得注意的是,该流匹配模型在采样时仅需5次函数求值(NFE),即可保持高增强质量,性能与使用约20次NFE时相当。实验表明,因果FM模型更适合少步数反向采样,且较小主干网络在长反向轨迹下性能下降。进一步对比显示,相同架构下,流匹配相比对抗损失训练更具优势。

原文摘要 · Abstract (English)

Generative models have shown robust performance on speech enhancement and restoration tasks, but most prior approaches operate offline with high latency, making them unsuitable for streaming applications. In this work, we investigate the feasibility of a low-latency, real-time generative speech restoration system based on flow-matching (FM). Our method tackles diverse real-world tasks, including denoising, dereverberation, and generative restoration. The proposed causal architecture without time-downsampling achieves introduces an total latency of only 20 ms, suitable for real-time communication. In addition, we explore a broad set of architectural variations and sampling strategies to ensure effective training and efficient inference. Notably, our flow-matching model maintains high enhancement quality with only 5 number of function evaluations (NFEs) during sampling, achieving similar performance as when using ~20 NFEs under the same conditions. Experimental results indicate that causal FM-based models favor few-step reverse sampling, and smaller backbones degrade with longer reverse trajectories. We further show a side-by-side comparison of FM to typical adversarial-loss-based training for the same model architecture.

语音修复流匹配实时生成低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。