arXiv:2601.13055eess.AS2026-01

VoCodec以极低计算量和30毫秒延迟实现高质量语音压缩,适合实时通信。

VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec

  • 采用轻量级结构与Vocos声码器结合,降低计算开销
  • 仅349.29M MACs/s、30ms延迟,在低码率语音编码挑战赛中排名第四
  • 可扩展语音增强功能,适合对实时性要求高的应用

端到端神经语音编解码器近年来可在极低比特率下保持高保真重建。同时,低计算复杂度与低延迟对实时通信至关重要。本文提出VoCodec,其计算复杂度仅为349.29M MACs/s,延迟为30毫秒。以竞争性声码器Vocos为骨干,该模型在2025年低码率音频编码挑战赛Track 1中排名第四,并在干净语音测试集上取得最高主观评分(MUSHRA)。此外,我们在前端级联轻量级神经网络以扩展语音增强能力。实验表明,两个系统在多个评估指标上表现优异。语音样本可访问https://acceleration123.github.io/。

原文摘要 · Abstract (English)

Recent advancements in end-to-end neural speech codecs enable compressing audio at extremely low bitrates while maintaining high-fidelity reconstruction. Meanwhile, low computational complexity and low latency are crucial for real-time communication. In this paper, we propose VoCodec, a speech codec model featuring a computational complexity of only 349.29M multiply-accumulate operations per second (MACs/s) and a latency of 30 ms. With the competitive vocoder Vocos as its backbone, the proposed model ranked fourth on Track 1 in the 2025 LRAC Challenge and achieved the highest subjective evaluation score (MUSHRA) on the clean speech test set. Additionally, we cascade a lightweight neural network at the front end to extend its capability of speech enhancement. Experimental results demonstrate that the two systems achieve competitive performance across multiple evaluation metrics. Speech samples can be found at https://acceleration123.github.io/.

语音编解码低延迟轻量化实时通信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。