arXiv:2509.17006cs.SDeess.AS2025-09被引 3

MBCodec通过分层解耦实现2.2kbps超高压缩音频还原。

MBCodec:Thorough disentangle for high-fidelity audio compression

  • 采用多码本与自监督语义分块,分离声学与语义特征
  • 实现170倍压缩率,24kHz音频仅需2.2kbps
  • 适合语音合成、低带宽传输等高保真场景

高保真神经语音编码器在文本转语音任务中旨在将语音信号压缩为离散表示以实现精准重建。然而,以往方法在有效解耦令牌中的声学与语义信息方面存在挑战,导致合成语音缺乏细粒度细节。本文提出MBCodec,一种基于残差向量量化(RVQ)的新型多码本音频编码器,学习分层结构化表示。该方法利用自监督语义分块和原始信号的子带特征构建功能解耦的潜在空间。为促进编码器各层嵌入空间的全面学习,引入自适应丢弃深度以差异化训练不同层的码本,并在训练中采用多通道伪正交镜像滤波器(PQMF)。通过彻底解耦语义与声学特征,本方法不仅实现近无损语音重建,还实现24kHz音频170倍压缩,比特率低至2.2kbps。实验验证其在各项评估中持续显著优于基线模型。

原文摘要 · Abstract (English)

High-fidelity neural audio codecs in Text-to-speech (TTS) aim to compress speech signals into discrete representations for faithful reconstruction. However, prior approaches faced challenges in effectively disentangling acoustic and semantic information within tokens, leading to a lack of fine-grained details in synthesized speech. In this study, we propose MBCodec, a novel multi-codebook audio codec based on Residual Vector Quantization (RVQ) that learns a hierarchically structured representation. MBCodec leverages self-supervised semantic tokenization and audio subband features from the raw signals to construct a functionally-disentangled latent space. In order to encourage comprehensive learning across various layers of the codec embedding space, we introduce adaptive dropout depths to differentially train codebooks across layers, and employ a multi-channel pseudo-quadrature mirror filter (PQMF) during training. By thoroughly decoupling semantic and acoustic features, our method not only achieves near-lossless speech reconstruction but also enables a remarkable 170x compression of 24 kHz audio, resulting in a low bit rate of just 2.2 kbps. Experimental evaluations confirm its consistent and substantial outperformance of baselines across all evaluations.

音频压缩语音生成分层编码低比特率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。