arXiv:2509.09550cs.SDcs.LG2025-09被引 3

用有限标量量化实现低比特率下抗干扰的神经音频压缩

Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates

  • 采用有限标量量化替代传统向量量化,简化训练并支持单一码本
  • 相同解码器下不同编码器可生成差异大但重建质量相当的码流
  • 在噪声信道模拟中,该方法比向量量化更抗比特错误,适合低带宽场景

神经音频编解码器(NACs)因优异的码率-失真表现及与大语言模型兼容的离散特征表示,在语音处理中日益普及。尽管多数现有编解码器依赖残差向量量化(RVQ),有限标量量化(FSQ)作为新兴替代方案,简化了训练过程并原生支持单码本。本文提出基于FSQ的NeuCodec,并证明其编码天然具备冗余性,能有效抵御传输噪声。首先,通过编码器蒸馏实验发现,两个不同编码器可将相同音频映射为差异显著的码序列,但使用相同量化器和解码器时重建质量相近。其次,模拟信道噪声传输时,FSQ编解码器在比特级扰动下的鲁棒性显著优于RVQ,尤其在低比特率下表现突出。

原文摘要 · Abstract (English)

Neural Audio Codecs (NACs) have become increasingly adopted in speech processing tasks due to their excellent rate-distortion performance and compatibility with Large Language Models (LLMs) as discrete feature representations for audio generation. While most existing codecs rely on Residual Vector Quantization (RVQ), Finite Scalar Quantization (FSQ) has recently emerged as a compelling alternative that simplifies training and natively supports single codebooks. We introduce NeuCodec, an FSQ-based NAC, and show that FSQ encodes baked-in redundancy which produces an encoding which is robust when transmitted through noisy channels. First, through an encoder distillation experiment, we show that two different encoders can learn to encode identical audio into vastly different code sequences whilst maintaining comparable reconstruction quality with the same quantizer and decoder. Second, we demonstrate that FSQ has vastly superior bit-level perturbation robustness by comparing the performance of RVQ and FSQ codecs when simulating the transmission of code sequences through a noisy channel.

音频压缩量化鲁棒传输低比特率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。