复杂谱建模提升音频编码鲁棒性,仅用30小时数据实现高保真跨域生成。
ComplexDec: A Domain-robust High-fidelity Neural Audio Codec with Complex Spectrum Modeling
- 采用复数频谱输入输出,减少压缩信息损失。
- 在24kbps比特率下,仅用30小时VCTK数据训练即达高保真效果。
- 适合追求低码率、跨域稳定性的音频生成应用。
神经音频编码器因其紧凑的离散表示,被广泛用于音频生成任务,但多数编码器在域外音频上表现不佳,导致下游生成任务误差传播。本文指出,编码压缩引发的信息丢失是造成域外鲁棒性下降的关键原因。为此,提出全频带48kHz的ComplexDec,采用复数频谱作为输入和输出,在保持与基线AudioDec和ScoreDec相同24kbps比特率的前提下,显著缓解信息损失。客观与主观评估表明,仅使用30小时VCTK语料库训练的ComplexDec,已具备优异的域外鲁棒性。
原文摘要 · Abstract (English)
Neural audio codecs have been widely adopted in audio-generative tasks because their compact and discrete representations are suitable for both large-language-model-style and regression-based generative models. However, most neural codecs struggle to model out-of-domain audio, resulting in error propagations to downstream generative tasks. In this paper, we first argue that information loss from codec compression degrades out-of-domain robustness. Then, we propose full-band 48~kHz ComplexDec with complex spectral input and output to ease the information loss while adopting the same 24~kbps bitrate as the baseline AuidoDec and ScoreDec. Objective and subjective evaluations demonstrate the out-of-domain robustness of ComplexDec trained using only the 30-hour VCTK corpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。