实现情绪可控且精准控制时长的零样本语音合成
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

- 提出新型时长控制方法,支持指定生成词数或自由生成
- 零样本下准确还原音色与指定情绪,情感保真度高
- 通过文本指令轻松控制情绪,适合影视配音等场景
现有自回归大规模语音合成模型在语音自然度上表现优异,但其逐词生成机制难以精确控制合成语音时长,这在视频配音等需严格音画同步的应用中成为显著瓶颈。本文提出IndexTTS2,引入一种通用且兼容自回归架构的时长控制新方法。该方法支持两种生成模式:一为显式指定生成词数以精确控制语音时长;二为自由生成,无需指定词数,仍能忠实还原输入提示的韵律特征。同时,模型实现情感表达与说话人身份的解耦,可独立控制音色与情绪。在零样本设置下,模型能准确重建目标音色(来自音色提示)并完美还原指定情绪(来自风格提示)。为提升高度情绪化表达下的语音清晰度,引入GPT隐变量表示,并设计三阶段训练范式增强生成稳定性。此外,通过微调Qwen3构建基于文本描述的软指令机制,有效引导生成具有预期情绪倾向的语音。多数据集实验表明,IndexTTS2在词错误率、说话人相似度和情感保真度方面均优于当前最优零样本语音合成模型。音频样例见:https://index-tts.github.io/index-tts2.github.io/
原文摘要 · Abstract (English)
Existing autoregressive large-scale text-to-speech (TTS) models have advantages in speech naturalness, but their token-by-token generation mechanism makes it difficult to precisely control the duration of synthesized speech. This becomes a significant limitation in applications requiring strict audio-visual synchronization, such as video dubbing. This paper introduces IndexTTS2, which proposes a novel, general, and autoregressive model-friendly method for speech duration control. The method supports two generation modes: one explicitly specifies the number of generated tokens to precisely control speech duration; the other freely generates speech in an autoregressive manner without specifying the number of tokens, while faithfully reproducing the prosodic features of the input prompt. Furthermore, IndexTTS2 achieves disentanglement between emotional expression and speaker identity, enabling independent control over timbre and emotion. In the zero-shot setting, the model can accurately reconstruct the target timbre (from the timbre prompt) while perfectly reproducing the specified emotional tone (from the style prompt). To enhance speech clarity in highly emotional expressions, we incorporate GPT latent representations and design a novel three-stage training paradigm to improve the stability of the generated speech. Additionally, to lower the barrier for emotional control, we designed a soft instruction mechanism based on text descriptions by fine-tuning Qwen3, effectively guiding the generation of speech with the desired emotional orientation. Finally, experimental results on multiple datasets show that IndexTTS2 outperforms state-of-the-art zero-shot TTS models in terms of word error rate, speaker similarity, and emotional fidelity. Audio samples are available at: https://index-tts.github.io/index-tts2.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。