arXiv:2511.22687cs.SDeess.AS2025-11中稿 · ASRU2025

用预训练语音增强模型分阶段量化,提升低码率语音编码稳定性与质量。

PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning

  • 分阶段量化:先处理低熵去噪特征,再编码高熵残差
  • 在噪声训练下仍优于传统方法,重建质量与下游任务表现更优
  • 适合追求高稳定性和高质量语音压缩的研究者

神经语音编码器在低码率压缩中表现优异,但残差向量量化(RVQ)常因训练不稳定和分解无效而影响重建质量与效率。本文提出PURE Codec(Progressive Unfolding of Residual Entropy),利用预训练语音增强模型引导多阶段量化:首阶段重建低熵、去噪的语音嵌入,后续阶段编码高熵残差成分。该设计显著提升训练稳定性。实验表明,PURE在重建质量及基于语音语言模型的文语转换任务中持续优于传统RVQ编码器,尤其在噪声训练条件下表现更佳。

原文摘要 · Abstract (English)

Neural speech codecs have achieved strong performance in low-bitrate compression, but residual vector quantization (RVQ) often suffers from unstable training and ineffective decomposition, limiting reconstruction quality and efficiency. We propose PURE Codec (Progressive Unfolding of Residual Entropy), a novel framework that guides multi-stage quantization using a pre-trained speech enhancement model. The first quantization stage reconstructs low-entropy, denoised speech embeddings, while subsequent stages encode residual high-entropy components. This design improves training stability significantly. Experiments demonstrate that PURE consistently outperforms conventional RVQ-based codecs in reconstruction and downstream speech language model-based text-to-speech, particularly under noisy training conditions.

语音编码向量量化语音增强分阶段学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。