用偏好优化让语音合成更精准表达情绪细微差别。
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
- 通过偏好优化对比不同情绪,提升情感区分能力。
- 在LJ Speech和Emo-DB数据集上显著优于基线模型。
- 适合需要精细控制情绪表达的语音合成应用。
当前的情感文本转语音(TTS)模型主要采用监督学习,将文本与目标情绪映射为对应语音,但每对文本-语音仅对应单一情绪。这类模型仅学习正确的情绪输出,未能充分理解不同情绪间的细微差异,限制了对情绪层次的捕捉能力。本文提出可控的Emo-DPO方法,利用直接偏好优化(Direct Preference Optimization),通过对比优选情绪与次优情绪,实现对情绪细微差别的有效区分。我们采用情感感知的LLM-TTS神经架构,借助大语言模型的上下文学习与指令遵循能力,替代传统TTS结构。大量实验表明,所提方法在情感辨识度和自然度上均超越现有基线模型。
原文摘要 · Abstract (English)
Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending other emotion characteristics, which limits their capabilities of capturing the nuances between different emotions. We propose a controllable Emo-DPO approach, which employs direct preference optimization to differentiate subtle emotional nuances between emotions through optimizing towards preferred emotions over less preferred emotional ones. Instead of relying on traditional neural architectures used in existing emotional TTS models, we propose utilizing the emotion-aware LLM-TTS neural architecture to leverage LLMs' in-context learning and instruction-following capabilities. Comprehensive experiments confirm that our proposed method outperforms the existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。