让语音情绪转换更精准可控,支持独立调节说话人和情感
Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
- 用不同参考音频分别控制内容、说话人和情感,实现属性解耦
- 引入时序情感表示与显式语调建模,提升情感动态迁移效果
- 适合需要精细情感控制的语音合成场景,如虚拟角色配音
情感语音转换(EVC)旨在保留语言内容的同时改变语音的情感风格。在实际应用中,可控性——即通过独立参考音频分别控制说话人身份和情感风格——至关重要。然而,现有方法往往难以完全解耦这些属性,且缺乏对细微情感表达(如时间动态)的建模能力。我们提出 Maestro-EVC,一种可控制的 EVC 框架,通过分别使用参考音频有效解耦内容、说话人身份和情感属性。此外,我们引入时序情感表示与显式语调建模,并结合语调增强策略,以鲁棒地捕捉并转移目标情感的时间动态特征,即使在语调不匹配条件下也能保持性能。实验结果表明,Maestro-EVC 能生成高质量、可控制且情感丰富的语音。
原文摘要 · Abstract (English)
Emotional voice conversion (EVC) aims to modify the emotional style of speech while preserving its linguistic content. In practical EVC, controllability, the ability to independently control speaker identity and emotional style using distinct references, is crucial. However, existing methods often struggle to fully disentangle these attributes and lack the ability to model fine-grained emotional expressions such as temporal dynamics. We propose Maestro-EVC, a controllable EVC framework that enables independent control of content, speaker identity, and emotion by effectively disentangling each attribute from separate references. We further introduce a temporal emotion representation and an explicit prosody modeling with prosody augmentation to robustly capture and transfer the temporal dynamics of the target emotion, even under prosody-mismatched conditions. Experimental results confirm that Maestro-EVC achieves high-quality, controllable, and emotionally expressive speech synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。