通过分步偏好优化,让语音合成更精准表达情绪。
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
- 在去噪过程中分步优化情绪偏好,实现细粒度控制
- 相比现有方法,语音情感表现力和自然度均提升
- 适合需要精确情绪控制的语音合成应用
情感文本到语音旨在传递情感的同时保持可懂性和语调自然,但现有方法依赖粗粒度标签或代理分类器,仅接收话语级反馈。本文提出情绪感知的分步偏好优化(EASPO),一种后训练框架,可在扩散模型的中间去噪步骤中对齐细粒度情绪偏好。核心是时间条件的EASPM模型,用于评分噪声中间语音状态,并实现自动偏好对构建。EASPO通过优化生成过程以匹配这些分步偏好,实现可控的情绪塑造。实验表明,在表现力和自然度方面均优于现有方法。
原文摘要 · Abstract (English)
Emotional text-to-speech seeks to convey affect while preserving intelligibility and prosody, yet existing methods rely on coarse labels or proxy classifiers and receive only utterance-level feedback. We introduce Emotion-Aware Stepwise Preference Optimization (EASPO), a post-training framework that aligns diffusion TTS with fine-grained emotional preferences at intermediate denoising steps. Central to our approach is EASPM, a time-conditioned model that scores noisy intermediate speech states and enables automatic preference pair construction. EASPO optimizes generation to match these stepwise preferences, enabling controllable emotional shaping. Experiments show superior performance over existing methods in both expressiveness and naturalness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。