arXiv:2505.24291cs.SDcs.AI2025-05被引 9

分离语音内容与韵律,实现精准零样本语音转换

Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion

  • 分离语音内容与韵律信息,用离散韵律标记控制风格
  • 在未见说话人上实现高保真语音转换,韵律控制精度显著提升
  • 适合需要精细风格控制的语音合成应用

当前零样本语音转换系统可合成未见说话人的声音,但多数方法难以准确还原源说话人的语调风格或模仿目标说话人的独特表达方式,限制了转换的可控性。本文提出Discl-VC框架,从自监督语音表征中解耦内容与韵律信息,并通过流匹配变换器结合上下文学习合成目标说话人语音。为实现生成语音韵律的精确控制,引入掩码生成变换器,基于提示非自回归地预测离散韵律标记。实验表明,Discl-VC在零样本语音转换中表现更优,且对合成语音的韵律控制具有显著准确性。

原文摘要 · Abstract (English)

Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic the distinctive speaking style of the target speaker, thereby limiting the controllability of voice conversion. In this work, we propose Discl-VC, a novel voice conversion framework that disentangles content and prosody information from self-supervised speech representations and synthesizes the target speaker's voice through in-context learning with a flow matching transformer. To enable precise control over the prosody of generated speech, we introduce a mask generative transformer that predicts discrete prosody tokens in a non-autoregressive manner based on prompts. Experimental results demonstrate the superior performance of Discl-VC in zero-shot voice conversion and its remarkable accuracy in prosody control for synthesized speech.

语音转换风格控制离散表征零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。