arXiv:2603.20638eess.AS2026-03被引 1

OmniCodec实现跨音频类型的低帧率通用编码,兼顾音质与语义信息。

OmniCodec: Low Frame Rate Universal Audio Codec with Semantic-Acoustic Disentanglement

  • 分层多码本设计结合语义-声学解耦,利用预训练理解模型编码器。
  • 同等码率下比Mimi codec重建质量更优,且生成任务语义信息更丰富。
  • 适合语音、音乐、通用声音等多场景音频生成,开源可用。

大语言模型通过离散表示学习推动了音频生成的发展。然而,现有神经编码器大多聚焦于语音,强调重建保真度,忽视了跨语音、音乐和通用声音等多样化音频领域的统一低帧率建模。此外,高重建质量并不必然带来语义信息丰富的表示,限制了下游生成任务的效果。为此,我们提出OmniCodec,一种专为低帧率设计的通用神经音频编码器。它采用分层多码本结构,通过利用预训练理解模型的音频编码器实现语义-声学解耦,并引入自指导策略提升码本利用率与重建效果。实验表明,在相同码率下,相比Mimi codec,OmniCodec在重建质量上表现更优,同时提供更具语义信息的表示,显著提升下游生成任务性能。模型与代码将公开发布,演示页面可访问。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have advanced audio generation through discrete representation learning. However, most existing neural codecs focus on speech and emphasize reconstruction fidelity, overlooking unified low frame rate modeling across diverse audio domains, including speech, music, and general sound. Moreover, high reconstruction quality does not necessarily yield semantically informative representations, limiting effectiveness in downstream generation tasks. We propose OmniCodec, a universal neural audio codec tailored for low frame rate. It adopts a hierarchical multi-codebook design with semantic-acoustic decoupling by leveraging the audio encoder of the pre-trained understanding model, along with a self-guidance strategy to improve codebook utilization and reconstruction. Compared with the Mimi codec, experiments show that OmniCodec achieves outstanding performance at the same bitrate, delivering superior reconstruction quality while also providing more semantically informative representations that benefit downstream generation tasks. Our model and code will be open-sourced. Our demo page is available.

音频编码语义解耦低帧率通用模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。