arXiv:2601.03888cs.SDcs.AI2026-01被引 4

IndexTTS 2.5提速增质,支持多语言零样本情感语音合成。

IndexTTS 2.5 Technical Report

  • 降低语义编码帧率至25Hz,压缩序列长度,降低训练与推理成本。
  • 用Zipformer替代U-DiT架构,参数减少且梅尔谱图生成更快。
  • 提出三种跨语言建模策略,实现无目标语言数据的情感迁移。

在前期工作中,我们提出了IndexTTS 2,一个由基于Transformer的文本到语义(T2S)模块和非自回归语义到梅尔(S2M)模块组成的零样本神经文本转语音基础模型,实现了忠实的情感复现,并建立了首个自回归时长可控生成范式。在此基础上,本文提出IndexTTS 2.5,通过四项关键改进显著提升多语言覆盖范围、推理速度和整体合成质量:1)语义编码压缩:将语义编码器帧率从50 Hz降至25 Hz,序列长度减半,大幅降低训练与推理开销;2)架构升级:将S2M模块的U-DiT骨干替换为更高效的Zipformer架构,实现显著参数减少与更快的梅尔谱图生成;3)多语言扩展:提出边界感知对齐、词元级拼接和指令引导生成三种显式跨语言建模范式,确立了支持中文、英文、日文、西班牙语的零样本情感语音合成实用设计原则,即使在缺乏目标语言情感训练数据下仍能实现稳健情感迁移;4)强化学习优化:在T2S模块后训练阶段引入GRPO,提升发音准确性和自然度。实验表明,IndexTTS 2.5不仅支持更广的语言覆盖,还能在相同零样本设置下复现未见语言的情感韵律。该模型实现RTF提升2.28倍,同时保持与IndexTTS 2相当的字错误率(WER)和说话人相似性。

原文摘要 · Abstract (English)

In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-based Text-to-Semantic (T2S) module and a non-autoregressive Semantic-to-Mel (S2M) module, which together enable faithful emotion replication and establish the first autoregressive duration-controllable generative paradigm. Building upon this, we present IndexTTS 2.5, which significantly enhances multilingual coverage, inference speed, and overall synthesis quality through four key improvements: 1) Semantic Codec Compression: we reduce the semantic codec frame rate from 50 Hz to 25 Hz, halving sequence length and substantially lowering both training and inference costs; 2) Architectural Upgrade: we replace the U-DiT-based backbone of the S2M module with a more efficient Zipformer-based modeling architecture, achieving notable parameter reduction and faster mel-spectrogram generation; 3) Multilingual Extension: We propose three explicit cross-lingual modeling strategies, boundary-aware alignment, token-level concatenation, and instruction-guided generation, establishing practical design principles for zero-shot multilingual emotional TTS that supports Chinese, English, Japanese, and Spanish, and enables robust emotion transfer even without target-language emotional training data; 4) Reinforcement Learning Optimization: we apply GRPO in post-training of the T2S module, improving pronunciation accuracy and natrualness. Experiments show that IndexTTS 2.5 not only supports broader language coverage but also replicates emotional prosody in unseen languages under the same zero-shot setting. IndexTTS 2.5 achieves a 2.28 times improvement in RTF while maintaining comparable WER and speaker similarity to IndexTTS 2.

语音合成多语言零样本高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。