用相对对抗反馈提升语音合成质量,小模型胜过大模型。
RAF: Relativistic Adversarial Feedback For Universal Speech Synthesis
- 用自监督语音模型辅助判别器评估波形质量,引导生成器学习更优表征。
- 在多个数据集上,参数减少88%的模型仍超越原大模型的感知质量。
- 适合想用小模型实现高性能语音合成的研究者与开发者。
我们提出相对对抗反馈(RAF),一种新型GAN声码器训练目标,可提升域内保真度并增强对未见场景的泛化能力。尽管现代GAN声码器采用先进架构,其训练目标常无法促进可泛化的表征学习。RAF通过利用语音自监督学习模型协助判别器评估样本质量,促使生成器学习更丰富的表示。此外,采用真实与虚假波形的相对配对方式,改善了训练数据分布的建模。跨多个数据集的实验表明,基于GAN的声码器在客观和主观指标上均取得一致提升。尤为重要的是,参数仅占12%的RAF训练的BigVGAN-base,在感知质量上优于全尺寸的LSGAN训练的BigVGAN。对比研究进一步验证了RAF作为GAN声码器训练框架的有效性。
原文摘要 · Abstract (English)
We propose Relativistic Adversarial Feedback (RAF), a novel training objective for GAN vocoders that improves in-domain fidelity and generalization to unseen scenarios. Although modern GAN vocoders employ advanced architectures, their training objectives often fail to promote generalizable representations. RAF addresses this problem by leveraging speech self-supervised learning models to assist discriminators in evaluating sample quality, encouraging the generator to learn richer representations. Furthermore, we utilize relativistic pairing for real and fake waveforms to improve the modeling of the training data distribution. Experiments across multiple datasets show consistent gains in both objective and subjective metrics on GAN-based vocoders. Importantly, the RAF-trained BigVGAN-base outperforms the LSGAN-trained BigVGAN in perceptual quality using only 12\% of the parameters. Comparative studies further confirm the effectiveness of RAF as a training framework for GAN vocoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。