无需外部监督,自监督实现零样本语音转换,兼顾隐私与音色还原。
GenVC: Self-Supervised Zero-Shot Voice Conversion
- 通过自监督分离说话人身份与语言内容,不依赖外部标注数据。
- 在零样本转换中达到更高说话人相似度,自然度媲美顶尖方法。
- 自回归设计灵活对齐时序,适合语音匿名化场景。
当前多数零样本语音转换方法依赖外部监督组件(尤其是说话人编码器)进行训练。为探索无需此类依赖的替代方案,本文提出GenVC,一种通过自监督方式从语音信号中解耦说话人身份与语言内容的新框架。GenVC以语音分词器和基于Transformer的自回归语言模型为骨干,支持大规模训练,同时增强源说话人隐私保护与目标说话人克隆保真度。实验表明,GenVC在零样本转换中实现了显著更高的说话人相似度,自然度与领先方法相当。由于其自回归形式,该方法在时间对齐上具有灵活性,降低了源语音韵律和说话人特性的保留,因而特别适用于语音匿名化。
原文摘要 · Abstract (English)
Most current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this dependency, this paper introduces GenVC, a novel framework that disentangles speaker identity and linguistic content from speech signals in a self-supervised manner. GenVC leverages speech tokenizers and an autoregressive, Transformer-based language model as its backbone for speech generation. This design supports large-scale training while enhancing both source speaker privacy protection and target speaker cloning fidelity. Experimental results demonstrate that GenVC achieves notably higher speaker similarity, with naturalness on par with leading zero-shot approaches. Moreover, due to its autoregressive formulation, GenVC introduces flexibility in temporal alignment, reducing the preservation of source prosody and speaker-specific traits, and making it highly effective for voice anonymization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。