arXiv:2409.05377eess.AScs.SD2024-09被引 103

1.04kbps下实现超清语音,性能远超同类模型。

BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

  • 模型规模达159M参数,十倍于主流模型,增强表达能力。
  • 在1.04kbps下主观听感优于原始语音,客观指标媲美6倍码率模型。
  • 融合序列模型与向量量化,高效捕捉语音时序特征。

我们提出 BigCodec,一种低比特率神经语音编解码器。尽管近期神经语音编解码器进展显著,但在低比特率(约1 kbps)下性能急剧下降。除码率限制外,模型容量等因素也制约进一步提升。为此,我们将模型规模扩大至159M参数,超过主流编解码器的10倍(约10M)。同时,将序列模型融入传统卷积架构以更好捕捉时序依赖,并采用低维向量量化以提高编码利用率。综合客观与主观评估表明,BigCodec 在1.04 kbps下显著优于多个现有低比特率编解码器;其客观性能相当于主流编解码器在4-6倍码率下的表现,甚至在主观感知质量上优于真实语音(ground truth)。

原文摘要 · Abstract (English)

We present BigCodec, a low-bitrate neural speech codec. While recent neural speech codecs have shown impressive progress, their performance significantly deteriorates at low bitrates (around 1 kbps). Although a low bitrate inherently restricts performance, other factors, such as model capacity, also hinder further improvements. To address this problem, we scale up the model size to 159M parameters that is more than 10 times larger than popular codecs with about 10M parameters. Besides, we integrate sequential models into traditional convolutional architectures to better capture temporal dependency and adopt low-dimensional vector quantization to ensure a high code utilization. Comprehensive objective and subjective evaluations show that BigCodec, with a bitrate of 1.04 kbps, significantly outperforms several existing low-bitrate codecs. Furthermore, BigCodec achieves objective performance comparable to popular codecs operating at 4-6 times higher bitrates, and even delivers better subjective perceptual quality than the ground truth.

语音编解码低码率神经网络向量量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。