arXiv:2603.01476eess.ASeess.SP2026-03

通过熵引导分组提升极低码率语音编码的保真度与语义质量

Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec

  • 基于通道方差分配信息量,实现编码输出的均衡分组
  • 在极低码率下显著提升语音可懂度与感知质量
  • 适合通信场景的高保真低码率语音编码

神经音频编解码器(NAC)对重建高质量语音信号和生成下游语音语言模型所需的离散表示至关重要。然而,在极低码率约束下同时保证准确的语义建模和高保真重建仍具挑战。本文提出熵引导的分组残差向量量化(EG-GRVQ),保留语义分支以捕捉语言信息,并在声学分支中引入熵引导分组策略。假设通道激活近似服从高斯分布,各通道方差可作为其信息含量的合理代理。据此,将编码器输出划分成若干组,使每组承担相等的信息总量。这种均衡分配提升了码本效率并减少了冗余。在LibriTTS和VCTK数据集上训练后,该模型在极低码率条件下展现出更高的感知质量与可懂度指标,特别适用于通信导向的编解码场景。

原文摘要 · Abstract (English)

Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity reconstruction under ultra-low bitrate constraints remains challenging. We propose an entropy-guided group residual vector quantization (EG-GRVQ) for an ultra-low bitrate neural speech codec, which retains a semantic branch for linguistic information and incorporates an entropy-guided grouping strategy in the acoustic branch. Assuming that channel activations follow approximately Gaussian statistics, the variance of each channel can serve as a principled proxy for its information content. Based on this assumption, we partition the encoder output such that each group carries an equal share of the total information. This balanced allocation improves codebook efficiency and reduces redundancy. Trained on LibriTTS and VCTK, our model shows improvements in perceptual quality and intelligibility metrics under ultra-low bitrate conditions, with a focus on codec-level fidelity for communication-oriented scenarios.

语音编码向量量化低码率熵引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。