首个支持逐词时长与停顿精确控制的语音合成模型
MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

- 通过显式词级时长条件与高置信度监督实现逐词控制
- 在自然语音合成基础上,词级时长与停顿控制误差降低40%以上
- 适合导航、阅读辅助等需要精细语速调整的应用场景
现有文本转语音系统缺乏细粒度局部时序控制:多数方法仅支持句级时长或整体语速控制,无法实现词级精确时序调节。MAGIC-TTS 是首个具备显式词级内容时长与停顿控制能力的 TTS 模型。其核心在于词级时长显式条件建模、高置信度时长监督数据构建,以及抑制零值偏差的训练机制。在时序控制基准测试中,MAGIC-TTS 显著提升词级时长与停顿跟随精度。即使无控制输入,仍能保持高质量自然语音合成。进一步在场景化编辑基准(导航指引、引导朗读、无障碍代码朗读)中验证,模型可建立统一时序基线,并将编辑区域精准移向目标时序,平均偏差低。结果表明,细粒度可控性可有效融入高质量 TTS 系统,支持真实局部时序编辑应用。
原文摘要 · Abstract (English)
Fine-grained local timing control is still absent from modern text-to-speech systems: existing approaches typically provide only utterance-level duration or global speaking-rate control, while precise token-level timing manipulation remains unavailable. To the best of our knowledge, MAGIC-TTS is the first TTS model with explicit local timing control over token-level content duration and pause. MAGIC-TTS is enabled by explicit token-level duration conditioning, carefully prepared high-confidence duration supervision, and training mechanisms that correct zero-value bias and make the model robust to missing local controls. On our timing-control benchmark, MAGIC-TTS substantially improves token-level duration and pause following over spontaneous synthesis. Even when no timing control is provided, MAGIC-TTS maintains natural high-quality synthesis. We further evaluate practical local editing with a scenario-based benchmark covering navigation guidance, guided reading, and accessibility-oriented code reading. In this setting, MAGIC-TTS realizes a reproducible uniform-timing baseline and then moves the edited regions toward the requested local targets with low mean bias. These results show that explicit fine-grained controllability can be implemented effectively in a high-quality TTS system and can support realistic local timing-editing applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。