用合成语音数据实现任意说话人间直接音色转换,提升音质与保真度。
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion
- 用预训练多说话人TTS生成同语义异音色的语音对进行训练
- 在未见说话人和新语言上实现零样本转换,词错率降低16.35%
- 适合需要高保真、跨说话人音色迁移的应用场景
传统语音转换方法通常试图将说话人身份与语言信息分离为独立表征,再进行重构。然而,有效解耦这些因素仍具挑战性,常导致训练过程中的信息丢失。本文提出一种新方法,利用高质量预训练多说话人文本转语音(TTS)模型生成的合成语音数据。具体地,使用具有相同语言内容但说话人不同的语音对作为输入-输出对来训练语音转换模型,使模型能直接学习源语音与目标语音之间的映射关系,有效捕捉说话人特有特征的同时保留语言内容。此外,我们提出一种灵活的训练策略,适用于任意到任意的语音转换,能良好泛化至未见说话人及新语言,在零样本场景下表现更优。实验表明,该方法在词错率上相对降低16.35%,说话人余弦相似度提升5.91%,优于多个先进方法。语音转换样例可访问:https://oovc-emnlp-2025.github.io/
原文摘要 · Abstract (English)
Traditional voice conversion (VC) methods typically attempt to separate speaker identity and linguistic information into distinct representations, which are then combined to reconstruct the audio. However, effectively disentangling these factors remains challenging, often leading to information loss during training. In this paper, we propose a new approach that leverages synthetic speech data generated by a high-quality, pretrained multispeaker text-to-speech (TTS) model. Specifically, synthetic data pairs that share the same linguistic content but differ in speaker identity are used as input-output pairs to train the voice conversion model. This enables the model to learn a direct mapping between source and target voices, effectively capturing speaker-specific characteristics while preserving linguistic content. Additionally, we introduce a flexible training strategy for any-to-any voice conversion that generalizes well to unseen speakers and new languages, enhancing adaptability and performance in zero-shot scenarios. Our experiments show that our proposed method achieves a 16.35% relative reduction in word error rate and a 5.91% improvement in speaker cosine similarity, outperforming several state-of-the-art methods. Voice conversion samples can be accessed at: https://oovc-emnlp-2025.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。