arXiv:2603.14032eess.AS2026-03

用跳跃扩散统一建模语音结构与内容,提升合成质量与自然度。

Beyond Two-stage Diffusion TTS: Joint Structure and Content Refinement via Jump Diffusion

  • 通过离散跳跃建模时间结构,连续扩散优化频谱细节,统一生成过程。
  • 单次推断即达3.37%错误率,优于Grad-TTS的4.38%,且语音质量更高。
  • 可自适应插入自然停顿,适合处理非标准语速的语音合成任务。

基于扩散和流匹配的文本转语音模型面临离散时间结构与连续频谱建模之间的矛盾。两阶段模型在固定对齐上进行扩散,常退化为平均语调;单阶段模型避免显式时长但存在对齐不稳定问题。本文提出一种跳跃扩散框架,其中离散跳跃用于建模时间结构,连续扩散在单一过程中精细化频谱内容。即使在单次推断的退化形式下,该方法在LJSpeech数据集上实现了3.37%的词错误率(WER),优于Grad-TTS的4.38%,同时提升了UTMOSv2得分。完整的迭代版本UDD进一步支持自适应语调,能在分布外的慢速语音中自主插入自然停顿,而非均匀拉伸。音频样例可访问https://anonymousinterpseech.github.io/TTS_Demo/。

原文摘要 · Abstract (English)

Diffusion and flow matching TTS faces a tension between discrete temporal structure and continuous spectral modeling. Two-stage models diffuse on fixed alignments, often collapsing to mean prosody; single-stage models avoid explicit durations but suffer alignment instability. We propose a jump-diffusion framework where discrete jumps model temporal structure and continuous diffusion refines spectral content within one process. Even in its one-shot degenerate form, our framework achieves 3.37% WER vs. 4.38% for Grad-TTS with improved UTMOSv2 on LJSpeech. The full iterative UDD variant further enables adaptive prosody, autonomously inserting natural pauses in out-of-distribution slow speech rather than stretching uniformly. Audio samples are available at https://anonymousinterpseech.github.io/TTS_Demo/.

语音合成扩散模型跳扩散自适应停顿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。