用音频编码器重构音乐,实现44.1kHz高保真合成。
High-Fidelity Music Vocoder using Neural Audio Codecs
- 将梅尔谱图映射到DAC隐空间,再用微调的DAC解码器还原音频。
- 在多个客观指标和听觉测试中均达到音乐合成新高度。
- 兼具音乐与语音合成能力,可作通用声码器使用。
尽管神经声码器在高保真语音合成方面取得显著进展,但其在多声部音乐中的应用仍鲜有探索。本文提出DisCoder,一种基于生成对抗编码器-解码器架构的神经声码器,利用神经音频编码器(Descript Audio Codec, DAC)的隐空间来从梅尔谱图重建44.1 kHz高保真音频。该方法首先将梅尔谱图转换为与DAC隐空间对齐的低维表示,再通过微调的DAC解码器重构音频信号。DisCoder在多个客观指标和MUSHRA听觉测试中均达到音乐合成的最先进水平。此外,其在语音合成任务上也表现优异,展现出作为通用声码器的巨大潜力。
原文摘要 · Abstract (English)
While neural vocoders have made significant progress in high-fidelity speech synthesis, their application on polyphonic music has remained underexplored. In this work, we propose DisCoder, a neural vocoder that leverages a generative adversarial encoder-decoder architecture informed by a neural audio codec to reconstruct high-fidelity 44.1 kHz audio from mel spectrograms. Our approach first transforms the mel spectrogram into a lower-dimensional representation aligned with the Descript Audio Codec (DAC) latent space before reconstructing it to an audio signal using a fine-tuned DAC decoder. DisCoder achieves state-of-the-art performance in music synthesis on several objective metrics and in a MUSHRA listening study. Our approach also shows competitive performance in speech synthesis, highlighting its potential as a universal vocoder.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。