改进测试时的量化编码,提升音频合成质量。
Improving Test-Time Performance of RVQ-based Neural Codecs
- 设计新编码算法,优化测试时代码选择以降低误差。
- 实验表明该方法有效减少量化误差并提升音质。
- 适合关注音频编码细节优化的研究者与工程师。
残差向量量化(RVQ)技术在近期神经音频编解码器中起核心作用。由于各量化层级间的层次结构,这些模型能用少量码本生成高保真音频。本文提出一种新的编码算法,旨在提升测试时RVQ类神经编解码器的合成质量。首先指出传统方法生成的量化向量存在次优性,证明通过选择不同代码集可减轻量化误差。随后提出新编码算法,用于寻找实现更低量化误差的离散代码组合。将该方法应用于预训练模型,并通过多种指标评估其有效性。实验结果验证:该方法不仅降低量化误差,还显著改善音频合成质量。
原文摘要 · Abstract (English)
The residual vector quantization (RVQ) technique plays a central role in recent advances in neural audio codecs. These models effectively synthesize high-fidelity audio from a limited number of codes due to the hierarchical structure among quantization levels. In this paper, we propose an encoding algorithm to further enhance the synthesis quality of RVQ-based neural codecs at test-time. Firstly, we point out the suboptimal nature of quantized vectors generated by conventional methods. We demonstrate that quantization error can be mitigated by selecting a different set of codes. Subsequently, we present our encoding algorithm, designed to identify a set of discrete codes that achieve a lower quantization error. We then apply the proposed method to pre-trained models and evaluate its efficacy using diverse metrics. Our experimental findings validate that our method not only reduces quantization errors, but also improves synthesis quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。