跨语言情感语音合成新框架,分离情绪与音色更自然
Perturbation Self-Supervised Representations for Cross-Lingual Emotion TTS: Stage-Wise Modeling of Emotion and Speaker
- 两阶段设计:先提取情绪特征,再恢复目标音色
- 引入扰动自监督表示,提升情绪与音色解耦效果
- 适合需要跨语言情感语音生成的研究者
跨语言情感语音合成旨在生成一种语言的语音,同时保留另一语言说话人的情绪,并保持目标语音的音色。该任务面临情绪、音色与语言三者高度耦合的挑战。为此,我们提出EMM-TTS,一种基于扰动自监督学习(SSL)表示的两阶段跨语言情感语音合成框架。第一阶段显式和隐式编码韵律线索以捕捉情绪表现力;第二阶段从扰动的SSL表示中恢复音色。我们研究了不同说话人扰动策略(基频偏移与说话人匿名化)对情绪与音色解耦的影响。为增强说话人保真度与表达控制,引入说话人一致性损失(SCL)和说话人-情绪自适应层归一化(SEALN)模块。此外,发现结合显式声学特征(如基频F0、能量、时长)与预训练潜在特征可提升语音克隆性能。综合多指标评估(主观与客观)表明,EMM-TTS在自然度、情绪迁移能力与音色一致性方面均优于现有方法。
原文摘要 · Abstract (English)
Cross-lingual emotional text-to-speech (TTS) aims to produce speech in one language that captures the emotion of a speaker from another language while maintaining the target voice's timbre. This process of cross-lingual emotional speech synthesis presents a complex challenge, necessitating flexible control over emotion, timbre, and language. However, emotion and timbre are highly entangled in speech signals, making fine-grained control challenging. To address this issue, we propose EMM-TTS, a novel two-stage cross-lingual emotional speech synthesis framework based on perturbed self-supervised learning (SSL) representations. In the first stage, the model explicitly and implicitly encodes prosodic cues to capture emotional expressiveness, while the second stage restores the timbre from perturbed SSL representations. We further investigate the effect of different speaker perturbation strategies-formant shifting and speaker anonymization-on the disentanglement of emotion and timbre. To strengthen speaker preservation and expressive control, we introduce Speaker Consistency Loss (SCL) and Speaker-Emotion Adaptive Layer Normalization (SEALN) modules. Additionally, we find that incorporating explicit acoustic features (e.g., F0, energy, and duration) alongside pretrained latent features improves voice cloning performance. Comprehensive multi-metric evaluations, including both subjective and objective measures, demonstrate that EMM-TTS achieves superior naturalness, emotion transferability, and timbre consistency across languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。