DPI-TTS让语音生成更快更自然,还更好保留说话人风格。
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
- 按频率和帧逐步生成语音,更符合声音特性
- 训练速度提升近2倍,效果优于基线模型
- 适合追求高效高质语音合成的研究者
近年来,语音扩散模型发展迅速。除了广泛使用的U-Net架构外,基于Transformer的模型如扩散Transformer(DiT)也受到关注。然而,当前的DiT语音模型将梅尔频谱图视为普通图像,忽略了语音特有的声学属性。为解决这一问题,我们提出一种名为方向性区块交互语音合成(DPI-TTS)的方法,基于DiT实现快速训练且不损失精度。DPI-TTS采用从低频到高频、逐帧推进的渐进式推理策略,更贴合语音声学特性,显著提升生成语音的自然度。此外,我们引入细粒度风格时序建模方法,进一步增强说话人风格相似性。实验结果表明,该方法使训练速度提升近2倍,并显著优于基线模型。
原文摘要 · Abstract (English)
In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。