arXiv:2510.21209eess.AScs.SD2025-10中稿 · Interspeech 2025被引 4

轻量级音频编码器,4kbps下性能超越主流模型

SpecTokenizer: A Lightweight Streaming Codec in the Compressed Spectrum Domain

  • 在压缩频谱域用交替卷积与循环网络建模
  • 4kbps下计算量仅需20%、参数量仅10%
  • 适合资源受限的实时语音应用

神经音频编解码器(NACs)近年来在语音模型的音频压缩与表征中受到广泛关注。尽管主流NACs通常需要千兆级计算和百万级参数,但轻量级流式NACs的性能仍待探索。本文提出SpecTokenizer,一种在压缩频谱域运行的轻量级流式编解码器。其仅由交替的卷积与循环神经网络层构成,通过在压缩频谱域进行多尺度建模,实现了更高效率与更强表征能力。在4 kbps速率下,SpecTokenizer的性能可媲美或超越当前最先进轻量级架构的编解码器,同时仅需20%的计算量和10%的参数量,且在相近资源条件下显著优于同类方案。

原文摘要 · Abstract (English)

Neural Audio Codecs (NACs) have gained growing attention in recent years as technologies for audio compression and audio representation in speech language models. While mainstream NACs typically require G-level computation and M-level parameters, the performance of lightweight and streaming NACs remains underexplored. This paper proposes SpecTokenizer, a lightweight streaming codec that operates in the compressed spectral domain. Composed solely of alternating CNN and RNN layers, SpecTokenizer achieves greater efficiency and better representational capability through multi-scale modeling in the compressed spectrum domain. At 4 kbps, the proposed SpecTokenizer achieves comparable or superior performance compared to the codec with state-of-the-art lightweight architecture while requiring only 20% of the computation and 10% of the parameters. Furthermore, it significantly outperforms the codec when using similar computational and storage resources.

音频编码轻量化流式处理频谱建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。