arXiv:2504.13791cs.SDcs.AI2025-04被引 1

用多判别器与最优传输优化语音转换,提升生成语音自然度。

Collective Learning Mechanism based Optimal Transport Generative Adversarial Network for Non-parallel Voice Conversion

  • 引入多判别器集体学习机制,融合DCNN、ViT与Conformer
  • 结合最优传输损失,精准匹配源与目标语音分布
  • 在三个数据集上均超越现有模型,适合语音合成研究者

生成对抗网络(GAN)在图像合成中取得显著成功后,也在语音合成领域进步明显,凭借对抗学习精准适配目标数据分布的能力。然而,当前最先进的基于GAN的语音转换(VC)模型在真实语音与生成语音之间仍存在显著自然度差异。此外,多数现有模型采用单生成器-单判别器结构,而单生成器多判别器架构更有利于优化目标数据分布。为此,本文提出一种新型GAN模型——基于集体学习机制的最优传输生成对抗网络(CLOT-GAN),集成深度卷积神经网络(DCNN)、视觉变压器(ViT)和Conformer等多种判别器,通过集体学习机制理解梅尔频谱的共振峰分布特征。同时,引入最优传输(OT)损失,依据最优传输理论精确弥合源与目标数据分布之间的差距。在VCC 2018、VCTK和CMU-Arctic数据集上的实验验证表明,所提出的CLOT-GAN-VC模型在客观与主观评估中均优于现有语音转换模型。

原文摘要 · Abstract (English)

After demonstrating significant success in image synthesis, Generative Adversarial Network (GAN) models have likewise made significant progress in the field of speech synthesis, leveraging their capacity to adapt the precise distribution of target data through adversarial learning processes. Notably, in the realm of State-Of-The-Art (SOTA) GAN-based Voice Conversion (VC) models, there exists a substantial disparity in naturalness between real and GAN-generated speech samples. Furthermore, while many GAN models currently operate on a single generator discriminator learning approach, optimizing target data distribution is more effectively achievable through a single generator multi-discriminator learning scheme. Hence, this study introduces a novel GAN model named Collective Learning Mechanism-based Optimal Transport GAN (CLOT-GAN) model, incorporating multiple discriminators, including the Deep Convolutional Neural Network (DCNN) model, Vision Transformer (ViT), and conformer. The objective of integrating various discriminators lies in their ability to comprehend the formant distribution of mel-spectrograms, facilitated by a collective learning mechanism. Simultaneously, the inclusion of Optimal Transport (OT) loss aims to precisely bridge the gap between the source and target data distribution, employing the principles of OT theory. The experimental validation on VCC 2018, VCTK, and CMU-Arctic datasets confirms that the CLOT-GAN-VC model outperforms existing VC models in objective and subjective assessments.

语音转换生成对抗网络最优传输多判别器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。