让语音合成精准控制每个词的情绪和语速,无需特殊标注数据。
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
- 通过自训练框架实现词级情感与语速的精细控制。
- 在无句内情感标注数据下仍达到顶尖表现。
- 适合需要自然动态语音生成的应用场景。
尽管情感语音合成已取得显著进展,但现有研究大多局限于话语级情感表达,难以支持词级控制。实现词级表现力控制面临根本挑战,主要源于多情感转换建模复杂性以及缺乏捕捉句内情感与语调变化的标注数据集。本文提出WeSCon,首个无需依赖包含句内情感或语速变化标注数据的自训练框架,可在预训练零样本语音合成模型中实现词级情感与语速控制。方法引入过渡平滑策略与动态语速控制机制,通过多轮推理引导模型完成词级表现力合成;为进一步简化推理,结合动态情感注意力偏置机制并进行自训练微调,以端到端方式激活模型的词级表现力控制能力。实验结果表明,WeSCon有效克服数据稀缺问题,在词级情感表达控制上达到当前最优性能,同时保持原始模型强大的零样本合成能力。
原文摘要 · Abstract (English)
While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modeling multi-emotion transitions and the scarcity of annotated datasets that capture intra-sentence emotional and prosodic variation. In this paper, we propose WeSCon, the first self-training framework that enables word-level control of both emotion and speaking rate in a pretrained zero-shot TTS model, without relying on datasets containing intra-sentence emotion or speed transitions. Our method introduces a transition-smoothing strategy and a dynamic speed control mechanism to guide the pretrained TTS model in performing word-level expressive synthesis through a multi-round inference process. To further simplify the inference, we incorporate a dynamic emotional attention bias mechanism and fine-tune the model via self-training, thereby activating its ability for word-level expressive control in an end-to-end manner. Experimental results show that WeSCon effectively overcomes data scarcity, achieving state-of-the-art performance in word-level emotional expression control while preserving the strong zero-shot synthesis capabilities of the original TTS model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。