无需伪对齐的多模态语音合成,自然度与效率双提升
M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
- 用多模态扩散变换器实现文本与语音的稳定对齐
- 英文和中文词错误率低至1.36%和1.31%,性能领先
- 适合追求高保真零样本语音合成的研究者
非自回归(NAR)语音合成依赖文本与音频表示之间的长度对齐,限制了自然度与表现力。现有方法依赖时长建模或伪对齐策略,严重制约自然度与计算效率。本文提出M3-TTS,一种基于多模态扩散变换器(MM-DiT)架构的简洁高效NAR语音合成范式。M3-TTS采用联合扩散变换层实现跨模态对齐,在无需伪对齐的前提下,稳定实现变长文本-语音序列的单调对齐;单个扩散变换层进一步增强声学细节建模。框架集成一个梅尔-变分自编码器(mel-vae)编解码器,带来3倍训练加速。在Seed-TTS和AISHELL-3基准上的实验表明,M3-TTS以最低的词错误率(英文1.36%,中文1.31%)达到当前最优的NAR性能,同时保持优异的自然度评分。代码与演示将公开于https://wwwwxp.github.io/M3-TTS。
原文摘要 · Abstract (English)
Non-autoregressive (NAR) text-to-speech synthesis relies on length alignment between text sequences and audio representations, constraining naturalness and expressiveness. Existing methods depend on duration modeling or pseudo-alignment strategies that severely limit naturalness and computational efficiency. We propose M3-TTS, a concise and efficient NAR TTS paradigm based on multi-modal diffusion transformer (MM-DiT) architecture. M3-TTS employs joint diffusion transformer layers for cross-modal alignment, achieving stable monotonic alignment between variable-length text-speech sequences without pseudo-alignment requirements. Single diffusion transformer layers further enhance acoustic detail modeling. The framework integrates a mel-vae codec that provides 3* training acceleration. Experimental results on Seed-TTS and AISHELL-3 benchmarks demonstrate that M3-TTS achieves state-of-the-art NAR performance with the lowest word error rates (1.36\% English, 1.31\% Chinese) while maintaining competitive naturalness scores. Code and demos will be available at https://wwwwxp.github.io/M3-TTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。