arXiv:2512.19442eess.SPcs.LG2025-12被引 6

32毫秒延迟的流式语音修复模型,实现实时通信中的高质量生成语音处理。

Real-Time Streamable Generative Speech Restoration with Flow Matching

  • 基于流形匹配的因果帧结构,支持低延迟生成
  • 总延迟仅48毫秒,24毫秒版本用于语音增强任务
  • 可在消费级显卡上运行,适用于多类语音任务

近年来,基于扩散的生成模型在语音处理领域产生了重大影响,展现出高自然度语音并催生了新研究方向。然而,由于其依赖多个大型深度神经网络的反复调用,计算开销大,限制了其在实时通信中的应用。本文提出 Stream.FM,一种基于流形匹配的因果帧生成模型,算法延迟为32毫秒,总延迟为48毫秒,为实时通信中的生成式语音处理铺平道路。我们设计了缓冲流式推理方案与优化的DNN架构,展示了学习到的少步数值求解器可在固定算力预算下提升输出质量,探索了模型权重压缩以找到算力与质量之间的平衡点,并提出一个24毫秒总延迟的变体用于语音增强任务。我们的工作超越理论延迟,证明了高质量流式生成语音处理可在当前消费级GPU上实现。Stream.FM可流式解决多种语音处理任务:语音增强、去混响、编解码后处理、带宽扩展、STFT相位恢复和梅尔声码合成。通过全面评估与MUSHRA听觉测试验证,Stream.FM在生成式流式语音修复方面达到最新水平,与非流式变体相比仅略有质量下降,且在生成式流式语音增强任务中优于近期工作Diffusion Buffer,同时延迟更低。

原文摘要 · Abstract (English)

Diffusion-based generative models have greatly impacted the speech processing field in recent years, exhibiting high speech naturalness and spawning a new research direction. Their application in real-time communication is, however, still lagging behind due to their computation-heavy nature involving multiple calls of large DNNs. Here, we present Stream$.$FM, a frame-causal flow-based generative model with an algorithmic latency of 32 milliseconds (ms) and a total latency of 48 ms, paving the way for generative speech processing in real-time communication. We propose a buffered streaming inference scheme and an optimized DNN architecture, show how learned few-step numerical solvers can boost output quality at a fixed compute budget, explore model weight compression to find favorable points along a compute/quality tradeoff, and contribute a model variant with 24 ms total latency for the speech enhancement task. Our work looks beyond theoretical latencies, showing that high-quality streaming generative speech processing can be realized on consumer GPUs available today. Stream$.$FM can solve a variety of speech processing tasks in a streaming fashion: speech enhancement, dereverberation, codec post-filtering, bandwidth extension, STFT phase retrieval, and Mel vocoding. As we verify through comprehensive evaluations and a MUSHRA listening test, Stream$.$FM establishes a state-of-the-art for generative streaming speech restoration, exhibits only a reasonable reduction in quality compared to a non-streaming variant, and outperforms our recent work (Diffusion Buffer) on generative streaming speech enhancement while operating at a lower latency.

语音修复流式生成扩散模型低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。