无需GAN也能生成高保真语音,训练更快更简单。
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
- 用幅度-相位串行预测替代GAN,简化模型结构。
- 语音质量接近顶尖GAN模型,训练速度提升显著。
- 适合追求高效训练的语音合成研究者与开发者。
近期主流基于梅尔谱图的神经声码器依赖生成对抗网络(GAN)实现高保真语音生成,如HiFi-GAN和BigVGAN。然而,GAN限制了训练效率与模型复杂度。本文提出新型FreeGAN声码器,旨在回答GAN是否为必要组件。FreeGAN采用幅度-相位串行预测框架,无需GAN训练,引入幅度先验输入、SNAKE-ConvNeXt v2主干网络及频率加权反缠绕相位损失,以弥补无GAN带来的性能下降。实验表明,FreeGAN的语音质量可媲美先进GAN基声码器,同时显著提升训练效率与模型复杂度。其他显式相位预测类神经声码器亦可借助本方法脱离GAN运行。
原文摘要 · Abstract (English)
Recently, mainstream mel-spectrogram-based neural vocoders rely on generative adversarial network (GAN) for high-fidelity speech generation, e.g., HiFi-GAN and BigVGAN. However, the use of GAN restricts training efficiency and model complexity. Therefore, this paper proposes a novel FreeGAN vocoder, aiming to answer the question of whether GAN is necessary for mel-spectrogram-based neural vocoders. The FreeGAN employs an amplitude-phase serial prediction framework, eliminating the need for GAN training. It incorporates amplitude prior input, SNAKE-ConvNeXt v2 backbone and frequency-weighted anti-wrapping phase loss to compensate for the performance loss caused by the absence of GAN. Experimental results confirm that the speech quality of FreeGAN is comparable to that of advanced GAN-based vocoders, while significantly improving training efficiency and complexity. Other explicit-phase-prediction-based neural vocoders can also work without GAN, leveraging our proposed methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。