WaveFM用流匹配实现高保真语音合成,一步生成音频且速度更快。
WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching
- 改用梅尔谱条件先验,降低语音合成中的能量传输成本。
- 引入多分辨率STFT辅助损失,显著提升音频质量。
- 定制一致性蒸馏加速推理,单步生成仍保持高保真度。
流匹配为扩散模型训练提供了稳定方法,但直接应用于神经声码器时效果不佳。本文提出WaveFM,一种基于流匹配的梅尔谱条件语音合成模型,旨在提升扩散声码器的样本质量和生成效率。由于梅尔谱表征波形的能量分布,WaveFM采用梅尔谱条件先验而非标准高斯先验,以减少合成过程中的无效能量转移。此外,不同于多数扩散声码器仅使用单一损失函数,我们引入改进的多分辨率短时傅里叶变换(STFT)辅助损失,进一步优化音质。为加速推理而不显著降低质量,提出针对WaveFM的定制一致性蒸馏方法。实验表明,相比以往扩散声码器,WaveFM在质量和效率上均表现更优,且可在单次前向传播中完成波形生成。
原文摘要 · Abstract (English)
Flow matching offers a robust and stable approach to training diffusion models. However, directly applying flow matching to neural vocoders can result in subpar audio quality. In this work, we present WaveFM, a reparameterized flow matching model for mel-spectrogram conditioned speech synthesis, designed to enhance both sample quality and generation speed for diffusion vocoders. Since mel-spectrograms represent the energy distribution of waveforms, WaveFM adopts a mel-conditioned prior distribution instead of a standard Gaussian prior to minimize unnecessary transportation costs during synthesis. Moreover, while most diffusion vocoders rely on a single loss function, we argue that incorporating auxiliary losses, including a refined multi-resolution STFT loss, can further improve audio quality. To speed up inference without degrading sample quality significantly, we introduce a tailored consistency distillation method for WaveFM. Experiment results demonstrate that our model achieves superior performance in both quality and efficiency compared to previous diffusion vocoders, while enabling waveform generation in a single inference step.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。