arXiv:2605.16251eess.AS2026-05

用低延迟架构让流匹配模型实时修复语音,算力降120倍

Real-time Speech Restoration using Data Prediction Mean Flows

  • 用数据预测均值流设计少步流匹配模型
  • 算力仅需现有方法的1/120,延迟几乎为零
  • 适合对延迟敏感的实时语音修复场景

生成模型能解决宽带扩展、音频间隙填补等非唯一解问题,以及去除编码器产生的高度非线性失真、削波和失真,而不仅是线性叠加噪声或混响。尽管大型离线模型已取得显著成果,但实时低延迟场景下的高性能模型仍待突破。本文提出一种结合数据预测均值流的少步流匹配模型,并搭配新颖的低延迟架构,使流匹配模型在实时约束下更具吸引力。相比当前最优方法,本模型算力降低120倍,除STFT外无额外算法延迟,同时保持相近的音频质量。

原文摘要 · Abstract (English)

Generative models are capable to address difficult problems with non-unique solutions like bandwidth extension and gap filling, removing highly non-linear artifacts from codecs, clipping and distortion, as opposed to removing linear additive components like noise and reverb. While large offline processing models have shown impressive results, these tasks have not been solved with real-time capable models with low latency and compute. We propose a few-step flow matching model using Data Prediction Mean Flows in combination with suitable novel low-latency architecture to make flow matching models an attractive choice under theses constraints. Compared to state-of-the-art, our proposed mean flow model uses 120x less compute and introduces no algorithmic latency other than the STFT, while achieving similar audio quality.

语音修复流匹配低延迟生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。