arXiv:2601.02776cs.SDcs.AI2026-01被引 2

单码本音频编码器实现低比特率高保真压缩,支持全频段高质量重建。

UniSRCodec: Unified and Low-Bitrate Single Codebook Codec with Sub-Band Reconstruction

  • 采用梅尔谱图时频联合压缩,结合声码器恢复相位信息。
  • 提出子带重建技术,实现高低频段均优质压缩,仅需40比特/秒。
  • 结构简洁超越现有单码本方法,适合语音通信与音频传输场景。

神经音频编码器(NACs)通过紧凑压缩与重建降低传输开销,弥合连续与离散信号的差距。现有NACs分为多码本与单码本两类:多码本结构复杂、下游适配难;单码本虽结构简单,却存在保真度低、统一建模能力弱、高频建模不足等问题。本文提出UniSRCodec,一种支持高采样率、低带宽、高保真与统一建模的单码本编码器。分析波形压缩效率低下,引入基于梅尔谱图的时频压缩方法,并配合声码器恢复原始音频相位。进一步提出子带重建技术,实现高低频段的高质量压缩。主观与客观实验表明,UniSRCodec在跨域单码本编码中达到最先进性能,仅需40比特/秒的码率,重建质量媲美部分多码本方法。

原文摘要 · Abstract (English)

Neural Audio Codecs (NACs) can reduce transmission overhead by performing compact compression and reconstruction, which also aim to bridge the gap between continuous and discrete signals. Existing NACs can be divided into two categories: multi-codebook and single-codebook codecs. Multi-codebook codecs face challenges such as structural complexity and difficulty in adapting to downstream tasks, while single-codebook codecs, though structurally simpler, suffer from low-fidelity, ineffective modeling of unified audio, and an inability to support modeling of high-frequency audio. We propose the UniSRCodec, a single-codebook codec capable of supporting high sampling rate, low-bandwidth, high fidelity, and unified. We analyze the inefficiency of waveform-based compression and introduce the time and frequency compression method using the Mel-spectrogram, and cooperate with a Vocoder to recover the phase information of the original audio. Moreover, we propose a sub-band reconstruction technique to achieve high-quality compression across both low and high frequency bands. Subjective and objective experimental results demonstrate that UniSRCodec achieves state-of-the-art (SOTA) performance among cross-domain single-codebook codecs with only a token rate of 40, and its reconstruction quality is comparable to that of certain multi-codebook methods. Our demo page is available at https://wxzyd123.github.io/unisrcodec.

音频编码单码本低比特率子带重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。