arXiv:2509.02244cs.SDcs.CL2025-09被引 2

用4x4频谱块量化替代复杂残差结构,实现低延迟高保真语音编码

Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding

  • 将梅尔频谱视为2D图像,直接对4x4块进行统一码本量化
  • 在7.5 kbits/s下达到与先进编码器相当的语音质量(STOI/PESQ)
  • 适合追求低延迟、可扩展性语音编码系统的研发者使用

我们提出一种新型神经语音编码器,通过引入单阶段非残差量化方法,挑战了传统复杂残差向量量化(RVQ)堆叠的必要性。该方法直接作用于梅尔频谱,将输入视为二维数据,对不重叠的4x4频谱块进行统一码本量化,构建离散潜在网格。为保证高保真合成,采用晚期对抗微调优化VQ-VAE,并从头训练一个HiFi-GAN声码器以还原编码后的频谱。系统在约7.5 kbits/s速率下运行,针对16 kHz语音进行评估,使用STOI、PESQ、MCD和ViSQOL等客观指标与多个前沿神经编码器对比。结果表明,该简化架构在感知质量与语音可懂度方面表现优异,验证其作为未来低延迟编码器设计的有效且开放基础。

原文摘要 · Abstract (English)

We present a neural speech codec that challenges the need for complex residual vector quantization (RVQ) stacks by introducing a simpler, single-stage quantization approach. Our method operates directly on the mel-spectrogram, treating it as a 2D data and quantizing non-overlapping 4x4 patches into a single, shared codebook. This patchwise design simplifies the architecture, enables low-latency streaming, and yields a discrete latent grid. To ensure high-fidelity synthesis, we employ a late-stage adversarial fine-tuning for the VQ-VAE and train a HiFi-GAN vocoder from scratch on the codec's reconstructed spectrograms. Operating at approximately 7.5 kbits/s for 16 kHz speech, our system was evaluated against several state-of-the-art neural codecs using objective metrics such as STOI, PESQ, MCD, and ViSQOL. The results demonstrate that our simplified, non-residual architecture achieves competitive perceptual quality and intelligibility, validating it as an effective and open foundation for future low-latency codec designs.

语音编码VQ-VAE低延迟声码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。