arXiv:2605.25506eess.AS2026-05

WaveNeXt 2用残差去噪和子模型提升语音合成速度与质量。

WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models

论文配图:WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models
图 1 · 摘自论文原文
  • 基于残差去噪和子模型逐步优化波形。
  • 扩散模型仅4步即可实现快速推理,训练仅32小时。
  • 兼容GAN与扩散模型,适合资源受限场景。

大多数神经声码器仅支持一种类型:生成对抗网络(GAN)或扩散模型。尽管最先进的模型如Vocos和WaveNeXt采用强大的ConvNeXt生成器,但仅用于GAN框架,在多说话人设置下性能有限。此外,扩散模型虽然训练比GAN快,但推理速度慢。本文提出WaveNeXt 2,一个兼容GAN与扩散模型的统一ConvNeXt框架。其核心创新为残差去噪与子模型设计,每个子模型逐步精炼波形。在多说话人数据集上的实验表明:(1) GAN-WaveNeXt 2比HiFi-GAN和WaveFit更快;(2) Diff-WaveNeXt 2在4步下推理极快且合成质量媲美FastDiff,同时训练仅需32小时,非常适合资源受限的应用。

原文摘要 · Abstract (English)

Most neural vocoders are limited to one type: either GAN or diffusion-based. While state-of-the-art models like Vocos and WaveNeXt use powerful ConvNeXt-based generators, they have only been used in GAN frameworks and have limited performance in multi-speaker settings. Moreover, diffusion models, despite training faster than GANs, have slow CPU inference. In this paper, we introduce WaveNeXt 2, a unified ConvNeXt-based framework compatible with both GAN and diffusion vocoders. Its core innovation is residual denoising and sub-modeling, where each sub-model progressively refines the waveform. Experimental results in the multi-speaker dataset demonstrate the effectiveness of our approach: (1) GAN-WaveNeXt 2 is much faster than HiFi-GAN and WaveFit, and (2) Diff-WaveNeXt 2 also delivers much faster inference and competitive synthesis quality compared with FastDiff with 4 steps. The Diff-WaveNeXt 2 is very training-efficient, training in only 32 hours, making it ideal for resource-constrained applications.

语音合成扩散模型ConvNeXt高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。