arXiv:2509.15085eess.AScs.LG2025-09被引 2

用流式生成流匹配实现低延迟语音合成,32毫秒算法延迟。

Real-Time Streaming Mel Vocoding with Generative Flow Matching

  • 基于生成流匹配与梅尔滤波器伪逆,设计可流式处理的语音声波生成模型。
  • 在16kHz语音上实现32毫秒算法延迟、48毫秒总延迟,实测可实时运行。
  • 相比非流式基线(如HiFi-GAN)显著提升PESQ和SI-SDR,音质更优。

Mel vocoding(即从梅尔谱图重建音频波形)仍是当前许多文本转语音系统的关键环节。基于生成流匹配、先前的生成STFT相位恢复方法(DiffPhase)以及梅尔滤波器组的伪逆算子,我们提出MelFlow——一种面向16 kHz语音的流式生成型梅尔声码器,具有仅32毫秒的算法延迟和48毫秒的总延迟。我们不仅在理论上证明了其实时流式能力,还在消费级笔记本显卡上实现了实际应用。此外,实验表明,该模型在PESQ和SI-SDR指标上显著优于多个非流式基线方法,包括HiFi-GAN。

原文摘要 · Abstract (English)

The task of Mel vocoding, i.e., the inversion of a Mel magnitude spectrogram to an audio waveform, is still a key component in many text-to-speech (TTS) systems today. Based on generative flow matching, our prior work on generative STFT phase retrieval (DiffPhase), and the pseudoinverse operator of the Mel filterbank, we develop MelFlow, a streaming-capable generative Mel vocoder for speech sampled at 16 kHz with an algorithmic latency of only 32 ms and a total latency of 48 ms. We show real-time streaming capability at this latency not only in theory, but in practice on a consumer laptop GPU. Furthermore, we show that our model achieves substantially better PESQ and SI-SDR values compared to well-established not streaming-capable baselines for Mel vocoding including HiFi-GAN.

语音合成流式生成流匹配声码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。