arXiv:2509.13670eess.AS2025-09中稿 · APSIPA ASC 2025

轻量级流式语音编码器,用知识蒸馏提升质量。

A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation

  • 全因果架构+通道压缩,实现低延迟流式编码。
  • 蒸馏高复杂度教师模型,重建质量媲美非流式方案。
  • 仅20毫秒延迟、910兆浮点运算、540万参数,适合实时通信。

当前许多神经语音编码器虽能实现高质量语音重建,但常忽视延迟与复杂度,限制其在实时语音通信和高效压缩等下游任务中的应用。此前我们提出StreamCodec,通过模型因果化和标量-向量组合量化策略实现可流式编码,但其重建质量与复杂度仍有提升空间。本文提出StreamCodec2,采用全因果架构并减少卷积通道数,实现轻量级流式编码。为弥补因因果化与剪枝导致的质量下降,引入非因果、高复杂度的教师编码器,通过知识蒸馏指导StreamCodec2训练。实验表明,采用知识蒸馏策略的StreamCodec2,在保持20毫秒低延迟、910 MFLOPs计算复杂度及5.4 M参数量的同时,实现了高质量语音重建。

原文摘要 · Abstract (English)

While many current neural speech codecs achieve impressive reconstructed speech quality, they often neglect latency and complexity considerations, limiting their practical deployment in downstream tasks such as real-time speech communication and efficient speech compression. In our previous work, we proposed StreamCodec, which enables streamable speech coding by leveraging model causalization and a scalar-vector-combined quantization strategy, but its reconstructed quality and complexity still have room for improvement. Therefore, this paper proposes an improved iteration of StreamCodec, named StreamCodec2. The StreamCodec2 supports streamable and lightweight speech coding by adopting a fully causal architecture and reducing the convolutional channels. To compensate for the speech quality degradation caused by model causalization and pruning, we introduce a non-causal, high-complexity teacher codec to guide the training of StreamCodec2 through knowledge distillation. Experimental results demonstrate that our proposed StreamCodec2, trained with the knowledge distillation strategy, can achieve high-quality speech reconstruction while maintaining low latency (only 20 ms), low computational complexity (only 910 MFLOPs), and low model complexity (only 5.4 M parameters).

语音编码流式处理知识蒸馏低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。