arXiv:2502.20067eess.AScs.SD2025-02ACL被引 31

统一音频编码器,用一个码本搞定语音、音乐和音效的高效压缩与生成。

UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook

  • 采用分域自适应码本与专家混合策略,统一处理多类音频数据。
  • 在语音、音乐、音效上均实现高保真重建,性能超越现有统一编码器。
  • 无额外模块下通过自监督掩码预测提升语义密度,适合大模型音频训练。

音频语言模型的发展依赖于神经音频编码器,其在连续波形与离散标记间建立映射,适配语言模型范式。从多层残差向量量化到单层量化的演进有利于语言自回归解码,但单一码本在跨领域音频(如语音、音乐、音效)上的表现受限于域间分布差异。本文提出UniCodec,一种基于单码本的统一音频编码器,支持多领域音频。通过分区域自适应码本方法与域混合专家策略,捕捉各音频领域的独特特征;此外,为在不引入额外模块的前提下增强编码器语义密度,提出自监督掩码预测建模方法。全面的客观与主观评估表明,UniCodec在三类音频上均实现优异重建性能,超越现有单码本统一编码器,并在声学与语义表征能力上超过最先进领域专用编码器。

原文摘要 · Abstract (English)

The emergence of audio language models is empowered by neural audio codecs, which establish critical mappings between continuous waveforms and discrete tokens compatible with language model paradigms. The evolutionary trends from multi-layer residual vector quantizer to single-layer quantizer are beneficial for language-autoregressive decoding. However, the capability to handle multi-domain audio signals through a single codebook remains constrained by inter-domain distribution discrepancies. In this work, we introduce UniCodec, a unified audio codec with a single codebook to support multi-domain audio data, including speech, music, and sound. To achieve this, we propose a partitioned domain-adaptive codebook method and domain Mixture-of-Experts strategy to capture the distinct characteristics of each audio domain. Furthermore, to enrich the semantic density of the codec without auxiliary modules, we propose a self-supervised mask prediction modeling approach. Comprehensive objective and subjective evaluations demonstrate that UniCodec achieves excellent audio reconstruction performance across the three audio domains, outperforming existing unified neural codecs with a single codebook, and even surpasses state-of-the-art domain-specific codecs on both acoustic and semantic representation capabilities.

音频编码统一建模自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。