轻量级音频编码器,4kbps下性能超越主流模型
SpecTokenizer: A Lightweight Streaming Codec in the Compressed Spectrum Domain
- 在压缩频谱域用交替卷积与循环网络建模
- 4kbps下计算量仅需20%、参数量仅10%
- 适合资源受限的实时语音应用
神经音频编解码器(NACs)近年来在语音模型的音频压缩与表征中受到广泛关注。尽管主流NACs通常需要千兆级计算和百万级参数,但轻量级流式NACs的性能仍待探索。本文提出SpecTokenizer,一种在压缩频谱域运行的轻量级流式编解码器。其仅由交替的卷积与循环神经网络层构成,通过在压缩频谱域进行多尺度建模,实现了更高效率与更强表征能力。在4 kbps速率下,SpecTokenizer的性能可媲美或超越当前最先进轻量级架构的编解码器,同时仅需20%的计算量和10%的参数量,且在相近资源条件下显著优于同类方案。
原文摘要 · Abstract (English)
Neural Audio Codecs (NACs) have gained growing attention in recent years as technologies for audio compression and audio representation in speech language models. While mainstream NACs typically require G-level computation and M-level parameters, the performance of lightweight and streaming NACs remains underexplored. This paper proposes SpecTokenizer, a lightweight streaming codec that operates in the compressed spectral domain. Composed solely of alternating CNN and RNN layers, SpecTokenizer achieves greater efficiency and better representational capability through multi-scale modeling in the compressed spectrum domain. At 4 kbps, the proposed SpecTokenizer achieves comparable or superior performance compared to the codec with state-of-the-art lightweight architecture while requiring only 20% of the computation and 10% of the parameters. Furthermore, it significantly outperforms the codec when using similar computational and storage resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。