通过情绪条件隐空间扩散,实现自然语音情感转换
TargetSEC: Plug-and-Play In-the-Wild Speech Emotion Conversion via Arousal-Conditioned Latent Style Diffusion
- 基于说话人与连续情绪条件生成风格嵌入,操作于紧凑隐空间
- 在MSP-Podcast上转换准确率更高,且无需时序建模仍保持高音质
- 适合需要即插即用、不依赖对齐数据的情感转换场景
语音情感转换(SEC)旨在将源语音的情绪转换为目标情绪,同时保持内容和说话人身份。在真实场景数据上的SEC极具挑战性,主要源于训练数据非平行以及复杂的实际声学环境。现有固定时长方法要么情感转换效果差(高质量但低转换率),要么降低语音自然度(低质量但高转换率)。本文提出TargetSEC,一种由嵌入驱动的隐空间扩散框架,可生成受说话人身份和连续情绪条件控制的情感聚焦风格嵌入。与在频谱图上扩散的方法不同,TargetSEC在紧凑的隐空间中运行。在MSP-Podcast数据集上的实验表明,TargetSEC在非时长基基线中实现了更高的转换准确率,同时保持了高语音质量,并且在不显式建模时间的情况下达到与时长预测系统相当的性能。
原文摘要 · Abstract (English)
Speech Emotion Conversion (SEC) aims to transform the emotion of a source utterance into a target emotion while preserving content and speaker identity. SEC on in-the-wild data is challenging due to the non-parallel nature of training data and complex real-world acoustics. Existing fixed-duration approaches either struggle to shift the emotion effectively (high quality, low conversion) or degrade speech naturalness (low quality, high conversion). We propose TargetSEC, an embedding-driven latent diffusion framework that generates emotion-focused style embeddings conditioned on speaker identity and continuous emotion. Unlike methods that diffuse over spectrograms, TargetSEC operates in a compact latent space. Experiments on the MSP-Podcast dataset show that TargetSEC outperforms current non-duration baselines in conversion accuracy while maintaining high speech quality, and achieves performance comparable to duration-prediction systems without explicit temporal modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。