提出新型神经音频压缩方法,实现低码率下高保真音质。
SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
- 采用稀疏量化扩展嵌入空间,提升压缩效率。
- 2.67 kbps下PESQ达2.87,音谱失真减少13%。
- 支持多码率且训练时间减半,适合实际部署。
神经音频压缩已成为高效表示语音、音乐和通用音频的有前景技术。然而,现有方法在低码率下性能显著下降,因可用嵌入空间严重受限。为此,我们提出一种通用高保真神经音频压缩算法,采用残差专家向量量化(REVQ),在带宽影响极小的情况下大幅扩展嵌入空间。引入温和的负载均衡策略,确保空间充分使用。此外,设计新型分层判别器,周期性地对STFT频谱进行分层,引导生成器关注关键频谱区域。为支持多码率且不损失低码率质量,采用高效的后训练策略。所提模型在2.67 kbps下取得2.87的PESQ和4.27的ViSQOL评分,有效降低频谱模糊,使与原始梅尔频谱图的距离减少13%。值得注意的是,后训练策略性能接近专用固定码率模型,同时将训练时间减少一半。大量消融实验验证了该方法优于基线。
原文摘要 · Abstract (English)
Neural audio compression has emerged as a promising technology for efficiently representing speech, music, and general audio. However, existing methods suffer from significant performance degradation at limited bitrates, where the available embedding space is sharply constrained. To address this, we propose a universal high-fidelity neural audio compression algorithm featuring Residual Experts Vector Quantization (REVQ), which substantially expands the embedding space with minimal impact on bandwidth. A gentle load-balancing strategy is introduced to ensure the full utilization of this expanded space. Furthermore, we develop a novel multi-tiered discriminator that periodically stratifies STFT spectra, guiding the generator to focus on critical spectral regions. To support multiple bitrates without quality loss at the lower end, we adopt an efficient post-training strategy. Our proposed model achieves impressive performance, with PESQ and ViSQOL scores of 2.87 and 4.27, respectively, at 2.67 kbps bandwidth. The approach effectively reduces spectral blur, decreasing the distance to the original mel-spectrogram by 13%. Notably, our post-training strategy achieves performance comparable to dedicated fixed-bitrate models while reducing the required training time by half. Extensive ablation studies confirm the superiority of our method over baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。