arXiv:2509.21968eess.AScs.SD2025-09被引 6

用单一码本实现语音与通用音频的高效压缩,700 bps下表现媲美专业编码器。

AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook

  • 单码本结合嵌套域特异性分区,通过教师模型蒸馏统一训练。
  • 16 kHz 混合域音频在约700 bps下实现高质量重建。
  • 适合需要通用音频压缩的场景,如多模态模型预训练。

我们提出AUV,一种基于单码本的统一神经音频编解码器,可高效重建语音并扩展至人声、音乐和声音等通用音频。AUV能在约700 bps的比特率下处理任意16 kHz混合域音频片段。通过嵌套的域特异性分区与对应教师模型引导的蒸馏机制,实现单阶段训练。采用以STFT特征为输入的Conformer风格编码器-解码器架构,提升了音频质量。全面评估表明,AUV在重建性能上可比肩当前最优的领域专用单层量化编解码器,展示了单码本实现音频通用向量量化的潜力。预训练模型与演示样本见https://swivid.github.io/AUV/。

原文摘要 · Abstract (English)

We propose AUV, a unified neural audio codec with a single codebook, which enables a favourable reconstruction of speech and further extends to general audio, including vocal, music, and sound. AUV is capable of tackling any 16 kHz mixed-domain audio segment at bit rates around 700 bps. To accomplish this, we guide the matryoshka codebook with nested domain-specific partitions, assigned with corresponding teacher models to perform distillation, all in a single-stage training. A conformer-style encoder-decoder architecture with STFT features as audio representation is employed, yielding better audio quality. Comprehensive evaluations demonstrate that AUV exhibits comparable audio reconstruction ability to state-of-the-art domain-specific single-layer quantizer codecs, showcasing the potential of audio universal vector quantization with a single codebook. The pre-trained model and demo samples are available at https://swivid.github.io/AUV/.

音频编码向量量化单码本通用音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。