解决语音编码中码本坍塌问题,提升音质与模型泛化能力。
ERVQ: Enhanced Residual Vector Quantization with Intra-and-Inter-Codebook Optimization for Neural Audio Codecs
- 通过码内平衡与码间多样性优化,防止码本使用不均。
- 在高保真语音编码中实现100%码本利用率,性能显著提升。
- 适用于各类神经语音编码器,尤其增强大模型语音生成效果。
当前神经语音编码器普遍采用残差向量量化(RVQ)对语音信号进行离散化,但常面临码本坍塌问题,导致有效码本规模下降,性能受限。为此,本文提出增强型残差向量量化(ERVQ),通过码内与码间双重优化机制缓解该问题。码内优化引入在线聚类与码本平衡损失,确保码本使用均衡高效;码间优化则通过降低连续量化结果间的相似性,提升特征多样性。实验表明,ERVQ在多种模型、采样率与比特率下均显著提升编码性能,尤其在最先进神经语音编码器上实现100%码本利用率。进一步实验显示,经ERVQ优化的编码器可有效提升统一语音-文本大语言模型的表现,下游零样本语音合成任务中生成语音自然度明显改善。
原文摘要 · Abstract (English)
Current neural audio codecs typically use residual vector quantization (RVQ) to discretize speech signals. However, they often experience codebook collapse, which reduces the effective codebook size and leads to suboptimal performance. To address this problem, we introduce ERVQ, Enhanced Residual Vector Quantization, a novel enhancement strategy for the RVQ framework in neural audio codecs. ERVQ mitigates codebook collapse and boosts codec performance through both intra- and inter-codebook optimization. Intra-codebook optimization incorporates an online clustering strategy and a code balancing loss to ensure balanced and efficient codebook utilization. Inter-codebook optimization improves the diversity of quantized features by minimizing the similarity between successive quantizations. Our experiments show that ERVQ significantly enhances audio codec performance across different models, sampling rates, and bitrates, achieving superior quality and generalization capabilities. It also achieves 100% codebook utilization on one of the most advanced neural audio codecs. Further experiments indicate that audio codecs improved by the ERVQ strategy can improve unified speech-and-text large language models (LLMs). Specifically, there is a notable improvement in the naturalness of generated speech in downstream zero-shot text-to-speech tasks. Audio samples are available here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。