MaskVCT实现零样本语音转换,支持多因素灵活控制。
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
- 通过多个无分类器引导机制整合多种条件,统一建模
- 在零样本设置下同时提升说话人相似度与语义可懂度
- 适合需要精细控制语音风格的语音合成与转换场景
我们提出MaskVCT,一种零样本语音转换(VC)模型,通过多重无分类器引导(CFGs)实现多因素可控性。不同于以往依赖固定条件设置的模型,MaskVCT在单一模型中集成多种条件。为增强鲁棒性和控制能力,该模型可选择使用连续或量化语言特征以提升可懂度和说话人相似性,也可选择是否使用音高轮廓来调控语调。这些灵活选项使用户能够在零样本语音转换场景中无缝平衡说话人身份、语言内容与语调因素。大量实验表明,MaskVCT在目标说话人相似度与口音相似度方面表现最优,同时在词错误率和字符错误率上达到与现有基线相当的水平。音频样例见 https://maskvct.github.io/。
原文摘要 · Abstract (English)
We introduce MaskVCT, a zero-shot voice conversion (VC) model that offers multi-factor controllability through multiple classifier-free guidances (CFGs). While previous VC models rely on a fixed conditioning scheme, MaskVCT integrates diverse conditions in a single model. To further enhance robustness and control, the model can leverage continuous or quantized linguistic features to enhance intelligibility and speaker similarity, and can use or omit pitch contour to control prosody. These choices allow users to seamlessly balance speaker identity, linguistic content, and prosodic factors in a zero-shot VC setting. Extensive experiments demonstrate that MaskVCT achieves the best target speaker and accent similarities while obtaining competitive word and character error rates compared to existing baselines. Audio samples are available at https://maskvct.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。