通过情感感知前缀实现语音转换中的显式情绪控制
Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
- 引入情感感知前缀,分阶段控制语音情绪特征
- 情绪转换准确率从42.40%提升至85.50%,保持说话人身份一致
- 适合需要精准情绪调控的语音合成与个性化语音应用
零样本语音转换近期在情绪控制方面展现出潜力,但因表达能力有限,性能不佳或不稳定。本文提出情感感知前缀(Emotion-Aware Prefix),应用于两阶段语音转换框架,显著提升情绪转换性能:情绪转换准确率(ECA)从42.40%提升至85.50%,同时保持语言内容完整性和语音质量,不破坏说话人身份。消融实验表明,序列调制与声学实现的联合控制对生成明显情绪至关重要。对比分析验证了方法的泛化能力,并揭示了声学解耦在维持说话人身份中的作用。
原文摘要 · Abstract (English)
Recent advances in zero-shot voice conversion have exhibited potential in emotion control, yet the performance is suboptimal or inconsistent due to their limited expressive capacity. We propose Emotion-Aware Prefix for explicit emotion control in a two-stage voice conversion backbone. We significantly improve emotion conversion performance, doubling the baseline Emotion Conversion Accuracy (ECA) from 42.40% to 85.50% while maintaining linguistic integrity and speech quality, without compromising speaker identity. Our ablation study suggests that a joint control of both sequence modulation and acoustic realization is essential to synthesize distinct emotions. Furthermore, comparative analysis verifies the generalizability of proposed method, while it provides insights on the role of acoustic decoupling in maintaining speaker identity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。