让语音情绪表达更细腻,通过排序优化实现精准情感强度控制
Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

- 将情感强度控制建模为列表偏好排序问题,学习文本与语音间的情感顺序关系
- 在固定文本下实现跨强度的连续情感表达,高情感强度下提升显著
- 构建多说话人情感强度数据集ESD-plus,支持精细化情感建模与评估
基于大语言模型(LLM)的文本转语音(TTS)系统虽能实现提示驱动的情感控制,但受限于文本与语音之间的语义-声学鸿沟,难以实现精细的情感强度调节。为此,本文将LLM-TTS中的情感强度控制建模为学习排序问题,提出Emo-LiPO框架,通过显式建模固定文本下各情感的全局强度排序,使语音生成更忠实、连续地反映文本中相对的情感强度。为进一步支持细粒度情感建模与评估,我们构建了多说话人数据集ESD-plus,其包含明确的情感强度变化。在ESD-plus上的实验表明,Emo-LiPO显著优于监督学习和基于DPO的基线方法,在情感准确性和强度可控性上均有提升,尤其在高强度情境下表现突出。
原文摘要 · Abstract (English)
Large language model (LLM)-based text-to-speech (TTS) systems enable prompt-conditioned emotional control but struggle with fine-grained emotion intensity due to the semantic -- acoustic gap between text and speech. To address this challenge, we formulate emotion intensity control in LLM-based TTS as a learning-to-rank problem and propose Emo-LiPO, a listwise preference optimization framework that aligns prompt-conditioned speech generation with relative emotion intensity expressed in text. Emo-LiPO explicitly models global intensity ordering within each emotion under fixed transcripts, enabling more faithful and continuous emotional expression. We further construct ESD-plus, a multi-speaker dataset with explicit emotion intensity variations, to support fine-grained emotion modeling and evaluation. Experiments on ESD-plus demonstrate that Emo-LiPO significantly improves emotion accuracy and intensity controllability over both supervised- and DPO-based LLM TTS baselines, with particularly pronounced gains at high intensity levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。