提出零样本情感语音转换新框架,分离情绪与音色,效果自然逼真。
ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism
- 用扩散模型解耦情绪与说话人特征,实现跨说话人情感迁移。
- 在未见说话人上达成高情感准确率与自然度,超越现有方法。
- 适合需要快速适配新说话人的情感语音应用开发者。
人类语音不仅传递语言内容,还蕴含情绪状态与个体特征。情感语音转换(EVC)旨在保留语义和说话人身份的同时改变情感表达,提升人机交互体验。尽管深度学习推动了特定目标说话人的EVC发展,现有方法仍存在情感准确性不足与语音失真等问题。此外,针对未见过的说话人进行零样本情感转换的研究尚不充分。本文提出一种基于扩散模型的解耦机制与表达引导的新框架,在大规模情感语音数据集上训练,并在域内与域外未见说话人数据集上评估。实验表明,该方法生成的语音具有高情感准确率、自然度与质量,展现出在更广泛EVC场景中的应用潜力。
原文摘要 · Abstract (English)
The human voice conveys not just words but also emotional states and individuality. Emotional voice conversion (EVC) modifies emotional expressions while preserving linguistic content and speaker identity, improving applications like human-machine interaction. While deep learning has advanced EVC models for specific target speakers on well-crafted emotional datasets, existing methods often face issues with emotion accuracy and speech distortion. In addition, the zero-shot scenario, in which emotion conversion is applied to unseen speakers, remains underexplored. This work introduces a novel diffusion framework with disentangled mechanisms and expressive guidance, trained on a large emotional speech dataset and evaluated on unseen speakers across in-domain and out-of-domain datasets. Experimental results show that our method produces expressive speech with high emotional accuracy, naturalness, and quality, showcasing its potential for broader EVC applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。