arXiv:2608.08432eess.AS2026-08

动态分配语音编码比特流,提升预训练模型压缩效率

BAMU: Bitstream-Aware Marginal-Utility Allocation for Frozen Pretrained Neural Speech Codecs

论文配图:BAMU: Bitstream-Aware Marginal-Utility Allocation for Frozen Pretrained Neural Speech Codecs
图 1 · 摘自论文原文
  • 根据帧和层的复杂度动态调整量化深度
  • 在固定比特率下实现3.45到3.78的主观评分提升
  • 适用于已冻结的预训练语音编码器,无需重新训练

预训练神经语音编码器通常对所有帧使用固定的残差向量量化(RVQ)深度,忽略了量化难度的时间变化。本文提出BAMU,一种面向冻结预训练编码器的比特流感知动态RVQ分配框架。一个轻量级、速率无关的预测器估计帧级与层级的边际潜在失真降低量,而约束分配器则在精确序列化大小预算下选择前缀有效深度。在LibriSpeech和VCTK数据集上的实验表明,对EnCodec有持续增益,对DAC在中高码率下表现更优。30人主观测试显示,主观评分为3.449提升至3.780,优于匹配的固定深度编码。

原文摘要 · Abstract (English)

Pretrained neural speech codecs typically use a fixed residual vector quantization (RVQ) depth for all frames, ignoring temporal variation in quantization difficulty. We propose BAMU, a bitstream-aware dynamic RVQ allocation framework for frozen pretrained codecs. A lightweight, rate-independent predictor estimates frame- and layer-wise marginal latent-distortion reductions, while a constrained allocator selects prefix-valid depths under an exact serialized-size budget. Experiments on EnCodec and DAC over LibriSpeech, together with VCTK evaluation, show consistent EnCodec gains and DAC improvements mainly at medium and high rates. A 30-listener study confirms a MOS improvement from 3.449 to 3.780 over matched fixed-depth coding.

语音编码量化分配预训练模型比特流优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。