用几何迭代检索提升神经音频编码器的音质重建效果
Geometric Iterative Retrieval for Neural Audio Codec Resynthesis
- 利用编码层层级结构在连续码本空间中迭代检索
- 在语音和音乐任务上优于单次预测与一步回归基线
- 适合需要高保真音频重建的研究者
基于残差向量量化(RVQ)的神经音频编码器已成为基于分词的一般音频生成主流离散表示,但从粗粒度编码令牌重建高质量音频仍是开放问题,限制了所有生成系统的保真度。先前工作将重建问题视为离散标记预测与连续回归之间的选择。我们认为这一二分法不完整,提出几何迭代检索,利用RVQ层级结构本身作为连续码本空间中的自然迭代分解。我们的方法不在离散词汇中分类或回归到单一目标向量,而是在码本的几何空间中进行对比检索。我们在语音和音乐的编码器恢复任务上评估该方法,结果表明其性能优于单次令牌预测和一步回归基线。
原文摘要 · Abstract (English)
Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every system that generates them. Prior work has framed resynthesis as a choice between discrete token prediction and continuous regression. We argue that this dichotomy is incomplete and introduce geometric iterative retrieval, a paradigm that uses the RVQ layer hierarchy itself as a natural iterative decomposition in continuous codebook space. Rather than classifying over discrete vocabularies or regressing to a single target vector, our method performs contrastive retrieval in the codebook's geometric space. We evaluate our method on codec restoration tasks across speech and music, and show improvements over both single-pass token prediction and one-step regression baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。