arXiv:2501.11999eess.AScs.SD2025-01被引 2

用新型熵模型替代传统量化器,提升语音压缩效率与质量。

Rate-Aware Learned Speech Compression

  • 用通道级熵模型替代量化器,简化训练并防止码本坍缩。
  • 在多种码率下实现53.51%的比特率节省,音质提升显著。
  • 适合需要高保真低码率语音传输的实时通信场景。

实时通信和大语言模型的快速发展,推动了语音压缩的重要性。基于深度学习的神经语音编解码器在率失真(RD)性能上已超越传统信号级编解码器。现有方法通常采用编码器-量化器-解码器架构,将音频转换为潜在特征表示后再转为离散符号。然而该架构存在两大缺陷:(1) 量化器性能不足,导致训练困难及码本坍缩问题;(2) 编码器与解码器表征能力有限,难以适配不同码率下的特征需求。本文提出一种率感知的自学习语音压缩方案,以先进的通道级熵模型替代量化器,从而提升率失真性能、简化训练流程并避免码本坍缩。同时引入多尺度卷积与线性注意力混合模块,增强编码器与解码器的表征能力与灵活性。实验表明,所提方法达到最先进的率失真性能,在平均情况下实现53.51%的BD-Rate比特率节省,并获得0.26的BD-VisQol与0.44的BD-PESQ增益。

原文摘要 · Abstract (English)

The rapid rise of real-time communication and large language models has significantly increased the importance of speech compression. Deep learning-based neural speech codecs have outperformed traditional signal-level speech codecs in terms of rate-distortion (RD) performance. Typically, these neural codecs employ an encoder-quantizer-decoder architecture, where audio is first converted into latent code feature representations and then into discrete tokens. However, this architecture exhibits insufficient RD performance due to two main drawbacks: (1) the inadequate performance of the quantizer, challenging training processes, and issues such as codebook collapse; (2) the limited representational capacity of the encoder and decoder, making it difficult to meet feature representation requirements across various bitrates. In this paper, we propose a rate-aware learned speech compression scheme that replaces the quantizer with an advanced channel-wise entropy model to improve RD performance, simplify training, and avoid codebook collapse. We employ multi-scale convolution and linear attention mixture blocks to enhance the representational capacity and flexibility of the encoder and decoder. Experimental results demonstrate that the proposed method achieves state-of-the-art RD performance, obtaining 53.51% BD-Rate bitrate saving in average, and achieves 0.26 BD-VisQol and 0.44 BD-PESQ gains.

语音压缩深度学习率失真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。