arXiv:2608.08362eess.AScs.HC2026-08

实现音素级语音表达控制,让合成语音更灵活自然。

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

论文配图:CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis
图 1 · 摘自论文原文
  • 分步控制:先全局说话人特征,再局部音素级韵律调节。
  • 零样本语音合成效果佳,且对语调、音量、时长可精细调节。
  • 适合需要精准控制语音情感与风格的场景,如配音、有声书。

近期文本转语音(TTS)系统在自然度和零样本语音克隆方面表现优异,但在词或音素级别进行细粒度表达控制仍具挑战。我们提出 CtrlSpeech,一种具有粗到精控制能力的可控表达式 TTS 框架。基于 DiTAR 架构,该方法结合全局说话人条件与音素对齐的音高、响度和持续时间信号,实现在保持目标说话人音色的前提下,对韵律进行局部精细调控。这一设计支持用户以细粒度时间分辨率调整语音表达属性,使语音优化更具灵活性与可控性。实验表明,CtrlSpeech 在零样本语音合成上表现优异,并显著提升对表达属性的控制能力,验证了其在灵活且实用的表达式语音合成中的有效性。

原文摘要 · Abstract (English)

Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme level remains challenging. We propose CtrlSpeech, a controllable, expressive TTS framework with coarse-to-fine control. Built on the DiTAR architecture, CtrlSpeech combines global speaker conditioning with phone-aligned pitch, loudness, and duration signals, enabling localized prosodic control while preserving the target speaker's timbre. This design allows users to adjust expressive attributes at a fine temporal granularity, making speech refinement more flexible and controllable. Experimental results show that CtrlSpeech achieves competitive zero-shot TTS performance and improves controllability over expressive attributes, demonstrating its effectiveness for flexible and practical expressive speech synthesis.

语音合成可控生成韵律控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。