arXiv:2506.16538cs.SDeess.AS2025-06中稿 · Interspeech 2025被引 4

动态分配码率,让语音编码在噪声环境下更省带宽且保质。

Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ

  • 按帧动态调整码率,优先保留语音关键成分
  • 噪声环境下比特率效率提升,主观听感更清晰
  • 适合实际场景中带噪语音的高效压缩

残差向量量化(RVQ)已成为神经语音与音频编码的主流方法,实现高保真压缩。然而,现实中的噪声会降低压缩效率,传统编码器均匀分配比特,浪费在不影响理解的噪声上。本文提出可变码率残差向量量化(VRVQ)框架,动态调整每帧码率以优化率失真权衡。相比固定码率(CBR)RVQ,该方法优先保留语音关键成分,抑制残差噪声,并集成特征去噪模块进一步增强抗噪能力。实验表明,VRVQ在噪声条件下优于传统方法,提升了压缩效率与感知质量。样本见项目页:https://yoongi43.github.io/noise_robust_vrvq/

原文摘要 · Abstract (English)

Residual Vector Quantization (RVQ) has become a dominant approach in neural speech and audio coding, providing high-fidelity compression. However, speech coding presents additional challenges due to real-world noise, which degrades compression efficiency. Standard codecs allocate bits uniformly, wasting bitrate on noise components that do not contribute to intelligibility. This paper introduces a Variable Bitrate RVQ (VRVQ) framework for noise-robust speech coding, dynamically adjusting bitrate per frame to optimize rate-distortion trade-offs. Unlike constant bitrate (CBR) RVQ, our method prioritizes critical speech components while suppressing residual noise. Additionally, we integrate a feature denoiser to further improve noise robustness. Experimental results show that VRVQ improves rate-distortion trade-offs over conventional methods, achieving better compression efficiency and perceptual quality in noisy conditions. Samples are available at our project page: https://yoongi43.github.io/noise_robust_vrvq/.

语音编码可变码率抗噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。