提升非自回归语音转换中的风格分离效果,减少源音色泄露。
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
- 用多语言离散语音单元表示内容,降低源音色干扰。
- 在情感和说话人相似度上优于基线模型,风格迁移更精准。
- 适合需要高质量情感语音转换的场景,如语音合成与角色配音。
表达性语音转换旨在将目标语音中的说话人身份和表达特征转移到源语音中。本文针对自监督、非自回归框架下的条件变分自编码器进行改进,重点减少源音色泄露并提升语言-声学特征的解耦能力。为降低风格泄露,采用多语言离散语音单元作为内容表示,并通过增强相似性损失和混风格层归一化强化嵌入表征。为增强表达性迁移,引入局部基频信息并通过交叉注意力机制建模,同时提取融合全局基频与能量特征的风格嵌入。实验表明,该模型在情感和说话人相似度指标上均优于基线,表现出更强的风格适应能力与更低的源风格泄露。
原文摘要 · Abstract (English)
Expressive voice conversion aims to transfer both speaker identity and expressive attributes from a target speech to a given source speech. In this work, we improve over a self-supervised, non-autoregressive framework with a conditional variational autoencoder, focusing on reducing source timbre leakage and improving linguistic-acoustic disentanglement for better style transfer. To minimize style leakage, we use multilingual discrete speech units for content representation and reinforce embeddings with augmentation-based similarity loss and mix-style layer normalization. To enhance expressivity transfer, we incorporate local F0 information via cross-attention and extract style embeddings enriched with global pitch and energy features. Experiments show our model outperforms baselines in emotion and speaker similarity, demonstrating superior style adaptation and reduced source style leakage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。