提出离散时间扩散模型,提升语音合成效率与一致性
Discrete-Time Diffusion-Like Models for Speech Synthesis
- 采用离散时间过程替代连续扩散,避免训练与推理不匹配
- 在相同质量下,推理步数显著减少,训练更高效
- 支持多种噪声类型,适合追求高效语音生成的研究者
扩散模型近年来备受关注,通常将语音生成建模为连续时间过程。为高效训练,该过程常受限于加性高斯噪声,具有局限性;推理时则需离散化时间,导致训练与采样条件不一致。相比之下,近期提出的离散时间过程无此类限制,可显著减少推理步数,且训练与推理完全一致。本文探索若干类扩散型离散时间过程,并提出新变体,包括加性高斯噪声、乘性高斯噪声、模糊噪声以及模糊与高斯噪声的混合。实验表明,这些离散时间过程在主观和客观语音质量上可媲美广泛使用的连续扩散模型,同时具备更高效、更一致的训练与推理机制。
原文摘要 · Abstract (English)
Diffusion models have attracted a lot of attention in recent years. These models view speech generation as a continuous-time process. For efficient training, this process is typically restricted to additive Gaussian noising, which is limiting. For inference, the time is typically discretized, leading to the mismatch between continuous training and discrete sampling conditions. Recently proposed discrete-time processes, on the other hand, usually do not have these limitations, may require substantially fewer inference steps, and are fully consistent between training/inference conditions. This paper explores some diffusion-like discrete-time processes and proposes some new variants. These include processes applying additive Gaussian noise, multiplicative Gaussian noise, blurring noise and a mixture of blurring and Gaussian noises. The experimental results suggest that discrete-time processes offer comparable subjective and objective speech quality to their widely popular continuous counterpart, with more efficient and consistent training and inference schemas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。