首个直接建模波形的零样本语音合成模型,实现端到端高质量语音生成。
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

- 直接建模原始波形,通过分块策略与多尺度频谱监督提升生成质量。
- 在公开数据集上接近现有最优隐空间模型性能,显著优于以往端到端模型。
- 首次证明波形空间扩散模型可规模化,为语音生成开辟新方向。
近期基于变分自编码器隐变量或梅尔频谱图的扩散模型已成为零样本语音合成的主流范式。尽管这些压缩表示提升了生成效率,但不可避免地造成信息损失且非端到端训练。理论上,直接建模原始波形可避免这些问题;然而该方向研究不足,常被认为因音频信号序列极长而困难。为此,我们提出 WavTTS,首个直接在波形空间建模的生成式零样本语音合成模型,显著缩小了与隐空间生成模型的差距。基于流匹配与扩散变换器(DiT),WavTTS采用简单的分块策略直接建模语音波形,同时引入多尺度梅尔频谱图监督以提供感知引导。此外,我们研究了波形扩散中的预测目标与噪声调度的影响,并设计了有效的调度方案以提升生成质量。在开源基准上的评估表明,WavTTS性能接近当前最先进的隐空间生成模型,显著优于此前的端到端语音生成模型。研究结果证实了在波形空间直接扩展基于扩散的语音合成的可行性,为端到端语音生成开辟了新路径。
原文摘要 · Abstract (English)
Recently, diffusion models operating on VAE latents or mel-spectrograms have become the dominant paradigm for zero-shot TTS. Although these compressed representations improve generation efficiency, they inevitably suffer from information loss and non-end-to-end training. Theoretically, directly modeling raw waveforms circumvents these issues; however, this direction remains underexplored and is often deemed difficult due to the extremely long sequence length of audio signals. To overcome this, we propose WavTTS, the first raw waveform generative TTS model that substantially narrows the gap with latent-space generative models. Built upon the flow matching with Diffusion Transformer (DiT), WavTTS directly models speech waveforms via a simple patchification strategy, while integrating multi-scale mel-spectrogram supervision to provide perceptual guidance during training. Furthermore, we investigate the impact of prediction targets and noise scheduling in waveform diffusion, and develop an effective schedule design to improve generation quality. Evaluations on open-source benchmarks demonstrate that WavTTS closely approaches the performance of current state-of-the-art latent generative zero-shot TTS models, while substantially outperforming previous end-to-end speech generation models. Our findings demonstrate the feasibility of scaling diffusion-based TTS directly in the waveform space, opening a new direction for end-to-end speech generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。