arXiv:2511.07135cs.SDeess.AS2025-11

无需目标语音即可生成全新真实人声,用于语音转换。

Generating Novel and Realistic Speakers for Voice Conversion

  • 用分层变分自编码器建模说话人音色空间
  • 采样生成的新说话人音色质量接近训练数据
  • 可插件式适配主流语音转换模型,无需重新训练

语音转换(VC)模型在保留副语言特征的同时改变音色,可用于配音和身份保护等场景。然而,多数VC系统需要目标语音数据,当目标数据不可用或用户希望转换为完全新颖、未见过的人声时,便受限。为此,我们提出轻量级方法SpeakerVAE,用于在VC中生成新说话人。该方法采用深层分层变分自编码器建模说话人音色空间,通过采样生成新型说话人表示,用于语音合成。所提方法可作为灵活插件模块,兼容多种VC模型,无需对基础模型进行联合训练或微调。我们在先进VC模型FACodec和CosyVoice2上验证了该方法,结果表明其能成功生成质量媲美训练数据的未见说话人。

原文摘要 · Abstract (English)

Voice conversion models modify timbre while preserving paralinguistic features, enabling applications like dubbing and identity protection. However, most VC systems require access to target utterances, limiting their use when target data is unavailable or when users desire conversion to entirely novel, unseen voices. To address this, we introduce a lightweight method SpeakerVAE to generate novel speakers for VC. Our approach uses a deep hierarchical variational autoencoder to model the speaker timbre space. By sampling from the trained model, we generate novel speaker representations for voice synthesis in a VC pipeline. The proposed method is a flexible plug-in module compatible with various VC models, without co-training or fine-tuning of the base VC system. We evaluated our approach with state-of-the-art VC models: FACodec and CosyVoice2. The results demonstrate that our method successfully generates novel, unseen speakers with quality comparable to that of the training speakers.

语音转换生成模型说话人建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。