纯风格化变压器+三元判别训练,实现高保真非平行语音转换
Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training
- 用分离式编码器与风格迁移解码器,解耦语音内容与音色
- 在多对多和多对一场景下,客观指标显著优于现有方法
- 适合语音克隆、数字人语音合成等需高质量音色迁移的场景
语音转换(VC)是智能人机交互的基础技术,旨在将任意源音色的语音转换为任意目标音色。基于生成对抗网络的传统方法在精确编码多样化语音元素并合成自然语音方面面临挑战。为此,我们提出Pureformer-VC,一种基于编码器-解码器架构的框架,采用Conformer块构建解耦编码器,使用Zipformer块设计风格迁移解码器。通过变分解耦训练(VAE)分离语音成分,并引入三元判别训练增强说话人区分能力。此外,结合注意力风格迁移机制(ASTM)与Zipformer共享权重,提升解码器风格迁移性能。在两个多说话人数据集上进行实验,结果表明,该模型在主观评价上与现有方法相当,但在客观指标上显著更优,尤其在多对多和多对一语音转换场景中表现突出。
原文摘要 · Abstract (English)
As a foundational technology for intelligent human-computer interaction, voice conversion (VC) seeks to transform speech from any source timbre into any target timbre. Traditional voice conversion methods based on Generative Adversarial Networks (GANs) encounter significant challenges in precisely encoding diverse speech elements and effectively synthesising these elements into natural-sounding converted speech. To overcome these limitations, we introduce Pureformer-VC, an encoder-decoder framework that utilizes Conformer blocks to build a disentangled encoder and employs Zipformer blocks to create a style transfer decoder. We adopt a variational decoupled training approach to isolate speech components using a Variational Autoencoder (VAE), complemented by triplet discriminative training to enhance the speaker's discriminative capabilities. Furthermore, we incorporate the Attention Style Transfer Mechanism (ASTM) with Zipformer's shared weights to improve the style transfer performance in the decoder. We conducted experiments on two multi-speaker datasets. The experimental results demonstrate that the proposed model achieves comparable subjective evaluation scores while significantly enhancing objective metrics compared to existing approaches in many-to-many and many-to-one VC scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。