提出分层表达向量,实现方言情感语音合成的零样本控制。
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
- 分两阶段构建方言与情感独立向量,再层级融合实现可控合成。
- 零样本下合成效果优于基线,方言自然度提升显著。
- 适合需要低成本适配方言情感的语音合成应用。
近期文本到语音(TTS)技术在自然度和可懂度上取得显著进展。研究逐渐转向提升语音表现力,如方言与情感合成。然而,结合方言与情感的跨风格合成仍具挑战,主要因缺乏带情感标签的方言数据。为此,我们提出分层表达向量(HE-Vector)方法,分两阶段实现情感化方言语音合成:第一阶段构建方言与情感独立任务向量,通过调节权重增强单风格合成,称为表达向量(E-Vector);第二阶段将这些向量分层融合,实现在无需联合标注数据下的可控情感表达方言合成。实验表明,HE-Vector在方言合成上表现优异,且在零样本设置下实现了有前景的情感表达方言语音生成。
原文摘要 · Abstract (English)
Recent advances in text-to-speech (TTS) have yielded remarkable improvements in naturalness and intelligibility. Building on these achievements, research has increasingly shifted toward enhancing the expressiveness of generated speech, such as dialectal and emotional TTS. However, cross-style synthesis combining both dialect and emotion remains challenging and largely unexplored, mainly due to the scarcity of dialectal data with emotional labels. To address this, we propose Hierarchical Expressive Vector (HE-Vector), a two-stage method for Emotional Dialectal TTS. In the first stage, we construct different task vectors to model dialectal and emotional styles independently, and then enhance single-style synthesis by adjusting their weights, a method we refer to as Expressive Vector (E-Vector). For the second stage, we hierarchically integrate these vectors to achieve controllable emotionally expressive dialect synthesis without requiring jointly labeled data, corresponding to Hierarchical Expressive Vector (HE-Vector). Experimental results demonstrate that HE-Vectors achieve superior performance in dialect synthesis, and promising results in synthesizing emotionally expressive dialectal speech in a zero-shot setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。