SwitchCodec动态切换专家量化器,实现高保真音频压缩。
Switchcodec: Adaptive residual-expert sparse quantization for high-fidelity neural audio coding
- 用动态路由的专家量化器替代固定码本,提升压缩效率。
- 在客观指标和主观听感测试中均优于现有方法。
- 支持推理时变比特率,无需重新训练即可多速率编码。
近期神经音频压缩模型常依赖残差向量量化实现高保真编码,但固定每帧码本数量对音频内容的广泛变化不具适应性,尤其对简单或复杂信号表现不佳。为此,我们提出SwitchCodec,基于残差专家向量量化(REVQ)的神经音频编解码器。REVQ结合共享量化器与动态路由的专家量化器,根据输入音频激活相应专家,解耦码本容量与比特率,确保各量化器全程可训练并充分使用。此外,可变比特率机制在推理时调整激活专家量化器数量,实现多比特率操作而无需重训练。实验表明,SwitchCodec在客观指标和主观听感测试中均超越现有基线。
原文摘要 · Abstract (English)
Recent neural audio compression models often rely on residual vector quantization for high-fidelity coding, but using a fixed number of per-frame codebooks is suboptimal for the wide variability of audio content-especially for signals that are either very simple or highly complex. To address this limitation, we propose SwitchCodec, a neural audio codec based on Residual Experts Vector Quantization (REVQ). REVQ combines a shared quantizer with dynamically routed expert quantizers that are activated according to the input audio, decoupling bitrate from codebook capacity and improving compression efficiency. This design ensures full training and utilization of each quantizer. In addition, a variable-bitrate mechanism adjusts the number of active expert quantizers at inference, enabling multi-bitrate operation without retraining. Experiments demonstrate that SwitchCodec surpasses existing baselines on both objective metrics and subjective listening tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。