用离散扩散模型替代自回归解码,实现更快更准的语音生成与识别。
Discrete Diffusion for Generative Modeling of Text-Aligned Speech Tokens
- 用离散扩散模型替代自回归语音解码器,提升生成效率。
- 10步去噪即可生成语音,单步生成质量下降微小。
- 在TASTE基础上优化量化模块,显著降低语音识别错误率。
本文提出一种用于文本对齐语音标记的离散扩散模型(DDM)框架。通过以离散扩散模型替代自回归语音解码器,该模型在语音重建质量、自动语音识别(ASR)性能和推理速度上均有显著提升。我们系统分析了将DDM应用于语音重建的多个方面,包括采样器选择、推理步数以及对长度估计误差的鲁棒性。此外,通过对比不同向量量化模块,发现FSQ相比RVQ可使自回归模型的相对词错误率(WER)降低35%,且提升0.14分通用语音质量评分(UT-MOS),同时增强扩散模型性能。所提模型仅需10步去噪即可生成语音,甚至支持单步生成,质量损失极小。
原文摘要 · Abstract (English)
This paper introduces a discrete diffusion model (DDM) framework for text-aligned speech tokenization and reconstruction. By replacing the auto-regressive speech decoder with a discrete diffusion counterpart, our model achieves significantly better reconstruction quality, stronger ASR performance, and faster inference. We provide a comprehensive analysis of applying DDMs to speech reconstruction, examining sampler choices, inference steps, and robustness to length-scale estimation errors. Furthermore, we improve the original TASTE by systematically comparing vector quantization modules, showing that FSQ yields up to a 35% relative WER reduction and +0.14 UT-MOS improvement over RVQ for AR models, while also enhancing DDM performance. Our model generates speech in just 10 denoising steps and even supports single-step generation with only minor quality degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。