arXiv:2512.20211cs.SDeess.AS2025-12中稿 · TASLP被引 1

解决神经音频合成中的混叠问题,提升音乐和人声合成质量。

Aliasing-Free Neural Audio Synthesis

  • 在激活函数和上采样模块中引入可微分抗混叠技术。
  • 在人声、音乐和音频上超越现有模型,语音表现相当。
  • 适用于高质量音乐与歌声生成,适合音频合成研究者。

在神经音频合成中,神经声码器和编解码器负责从声学或潜在表示重建波形,对音频质量至关重要。尽管当前模型能生成感知自然的语音,但在高保真音乐和歌唱声合成方面仍受非线性激活函数和上采样层引入严重混叠伪影的制约。虽然数字信号处理中已有多种抗混叠技术,但其在神经声码器和编解码器中的集成仍较少被探索。本文将可微分抗混叠技术融入激活和上采样模块,提出 Pupu-Vocoder 与 Pupu-Codec。我们构建了一个测试信号基准以评估抗混叠模块,并在语音、歌唱声、音乐和音频数据上验证了所提模型。实验结果表明,Pupu-Vocoder 与 Pupu-Codec 在歌唱声、音乐和音频任务上优于现有系统,而在语音任务上表现相当。演示、代码与检查点可在 VocodexElysium.github.io/AliasingFreeNeuralAudioSynthesis/ 获取。

原文摘要 · Abstract (English)

In neural audio synthesis, neural vocoders and codecs are models that reconstruct waveforms from acoustic and latent representations, which are essential to the resulting audio quality. While current models are capable of generating perceptually natural speech, they still struggle with high-fidelity music and singing voice synthesis, as severe aliasing artifacts are introduced by non-linear activation functions and upsampling layers in existing architectures. Although various anti-aliasing techniques have been proposed in digital signal processing, their integration into neural vocoders and codecs remains under-explored. This paper incorporates differentiable anti-aliasing techniques into the activation and upsampling modules to bridge this gap, and thus presents Pupu-Vocoder and Pupu-Codec. We build a test signal benchmark to evaluate the anti-aliased modules, and validate our proposed models on speech, singing voice, music, and audio. Experimental results show that Pupu-Vocoder and Pupu-Codec outperform existing systems on singing voice, music, and audio, while achieving comparable performance on speech. Demos, codes, and checkpoints are available at: VocodexElysium.github.io/AliasingFreeNeuralAudioSynthesis/.

音频合成抗混叠声码器音乐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。