250 bps下实现高保真语音编码,兼顾自然度与说话人特征。
Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction

- 三阶段框架:编码-精炼-波形重建,提升极端低比特率下的语音质量。
- 640倍压缩+动态聚类,防止码本退化,保持多样性。
- 轻量级流匹配精炼,仅需少量迭代即可高效还原语音细节。
超低比特率语音编码对带宽受限通信和深度压缩至关重要,但在极端比特预算下维持语音自然度与说话人身份仍具挑战,主要源于信息丢失严重和量化不稳定性。为此,我们提出FMelCodec,一种基于梅尔频谱域的超低比特率神经语音编解码器,采用三阶段编码-精炼-重建(CRR)框架,最低可支持250 bps。前端梅尔频谱编码阶段采用高度激进的640×压缩/解压编码器-解码器结构,配备单个1024个码字的向量量化(VQ)码本,并结合在线聚类策略重新分配未使用码字,防止码本崩溃并维持码本多样性。后续条件流匹配(CFM)精炼阶段利用轻量级速度场估计器和基于CFM的求解器,对前序解码器生成的退化梅尔频谱进行精炼,并采用自一致性训练方案,支持较少迭代推理步骤以降低计算开销。最终波形重建阶段采用HiFi-GAN声码器,从精炼后的梅尔频谱中忠实还原波形。在两个涵盖两种采样率的数据集上的实验表明,在16 kHz下250 bps、48 kHz下750 bps的超低比特率约束下,客观与主观评估均一致显示,FMelCodec在语音重建质量与说话人相似性方面表现更优,且计算与模型复杂度更低。
原文摘要 · Abstract (English)
Ultra-low-bitrate speech coding is pivotal for bandwidth-constrained communication and deep compression, yet maintaining naturalness and speaker identity at such extreme bit budgets remains challenging due to pronounced information loss and quantization instability. To this end, we propose FMelCodec, an ultra-low-bitrate neural speech codec in the mel-spectrogram domain, cast as a three-stage coding-refinement-reconstruction (CRR) framework that can operate at as low as 250 bps. In the CRR framework, the front-end mel-spectrogram coding stage employs a highly aggressive 640x compression/decompression encoder-decoder structure with a single 1024-entry VQ codebook, coupled with an online clustering strategy that reassigns underused codewords to prevent codebook collapse and preserve codebook diversity. The subsequent conditional flow matching (CFM)-based mel-spectrogram refinement stage leverages a lightweight velocity-field estimator and CFM-based solver to refine the codec-degraded mel-spectrogram produced by the preceding decoder, and adopts a self-consistency training scheme that supports fewer iterative inference steps for the purpose of reducing computational overhead. Finally, the vocoding-driven waveform reconstruction stage employs a HiFi-GAN vocoder to faithfully reconstruct waveform from the refined mel-spectrogram. Experiments conducted on two datasets spanning two sampling rates show that, under ultra-low-bitrate constraints of 250 bps for 16 kHz and 750 bps for 48 kHz, both objective and subjective evaluations consistently demonstrate that FMelCodec achieves higher speech reconstruction quality and speaker similarity, while incurring lower computational and model complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。