通过因果建模实现语音情感的精准控制,让同一文本可输出不同情绪。
Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2
- 构建因果模型,分离情感与语义对语音韵律的影响。
- 在多说话人数据集上提升情感识别准确率和语音自然度。
- 支持跨说话人情感迁移,适合需要可控表达的语音应用。
我们提出一种新的因果韵律中介框架,用于生成富有表现力的文本到语音(TTS)合成。该方法在FastSpeech2架构中引入显式情感条件,并设计反事实训练目标,以解耦情感韵律与语言内容。通过建立文本(内容)、情感和说话人共同影响韵律(时长、音高、能量)并最终生成语音波形的结构因果模型,推导出两项互补损失:间接路径约束(IPC),强制情感仅通过韵律影响语音;反事实韵律约束(CPC),鼓励不同情感对应不同的韵律模式。模型在多说话人情感语料库(LibriTTS、EmoV-DB、VCTK)上联合训练,包含标准频谱重建和方差预测损失以及我们的因果损失。评估显示,该方法在表达性语音合成中显著提升韵律操控能力与情感呈现效果,平均意见分(MOS)和情感准确率均高于基线。同时在跨说话人情感迁移中表现出更低的词错误率(WER)和更好的说话人一致性。消融实验验证了因果目标成功分离韵律归因,实现可解释的反事实韵律编辑(如“相同语句,不同情绪”),且不牺牲自然度。本文讨论了韵律建模中的可辨识性问题,并指出假设情感影响完全由音高、时长和能量捕捉的局限性。结果表明,将因果学习融入TTS可显著提升可控性与表现力。
原文摘要 · Abstract (English)
We propose a novel causal prosody mediation framework for expressive text-to-speech (TTS) synthesis. Our approach augments the FastSpeech2 architecture with explicit emotion conditioning and introduces counterfactual training objectives to disentangle emotional prosody from linguistic content. By formulating a structural causal model of how text (content), emotion, and speaker jointly influence prosody (duration, pitch, energy) and ultimately the speech waveform, we derive two complementary loss terms: an Indirect Path Constraint (IPC) to enforce that emotion affects speech only through prosody, and a Counterfactual Prosody Constraint (CPC) to encourage distinct prosody patterns for different emotions. The resulting model is trained on multi-speaker emotional corpora (LibriTTS, EmoV-DB, VCTK) with a combined objective that includes standard spectrogram reconstruction and variance prediction losses alongside our causal losses. In evaluations on expressive speech synthesis, our method achieves significantly improved prosody manipulation and emotion rendering, with higher mean opinion scores (MOS) and emotion accuracy than baseline FastSpeech2 variants. We also observe better intelligibility (low WER) and speaker consistency when transferring emotions across speakers. Extensive ablations confirm that the causal objectives successfully separate prosody attribution, yielding an interpretable model that allows controlled counterfactual prosody editing (e.g. "same utterance, different emotion") without compromising naturalness. We discuss the implications for identifiability in prosody modeling and outline limitations such as the assumption that emotion effects are fully captured by pitch, duration, and energy. Our work demonstrates how integrating causal learning principles into TTS can improve controllability and expressiveness in generated speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。