用方向向量调控情绪强度,让语音转换更精准自然。
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
- 通过无监督方向潜向量建模,动态调节情绪嵌入强度。
- 在英、印语种上超越现有方法,主观与客观评估均提升。
- 适合语音合成、情感计算领域研究者参考。
情感语音转换(EVC)旨在将语音中的情感状态从源情绪精确转换为目标情绪,同时保持语言内容不变。本文提出在基于扩散模型的EVC框架中,通过自监督特征表示与无监督方向潜向量建模(DVM)来规范情绪强度,实现高保真情感生成。传统方法依赖情绪类别概率或强度标签,常导致风格控制不准且音质下降。本文方法在情感嵌入空间中利用目标情绪强度和方向向量动态调整嵌入,并在逆扩散过程中融合更新后的嵌入以生成目标情感与强度的语音。该工作首次实现了扩散模型框架下的情绪强度规范化,已在英文与印地语上验证,优于当前最优基线,在主观与客观评价中均表现优异。
原文摘要 · Abstract (English)
The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regularizing emotion intensity in the diffusion-based EVC framework to generate precise speech of the target emotion. Traditional approaches control the intensity of an emotional state in the utterance via emotion class probabilities or intensity labels that often lead to inept style manipulations and degradations in quality. On the contrary, we aim to regulate emotion intensity using self-supervised learning-based feature representations and unsupervised directional latent vector modeling (DVM) in the emotional embedding space within a diffusion-based framework. These emotion embeddings can be modified based on the given target emotion intensity and the corresponding direction vector. Furthermore, the updated embeddings can be fused in the reverse diffusion process to generate the speech with the desired emotion and intensity. In summary, this paper aims to achieve high-quality emotional intensity regularization in the diffusion-based EVC framework, which is the first of its kind work. The effectiveness of the proposed method has been shown across state-of-the-art (SOTA) baselines in terms of subjective and objective evaluations for the English and Hindi languages \footnote{Demo samples are available at the following URL: \url{https://nirmesh-sony.github.io/EmoReg/}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。