超低码率音频编码器,专为语音大模型设计,兼顾清晰度与实时性。
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
- 分阶段训练+解耦架构,提升语义建模与特征提取能力。
- 帧率仅16.67 Hz,码率0.43~0.87 kbps,仍保高可懂性。
- 适合语音大模型部署,支持低延迟流式合成,工业级可用。
本文提出LongCat-Audio-Codec,一种面向工业级端到端语音大模型的音频分词器与反分词器方案。通过解耦模型架构与多阶段训练策略,该方案具备强语义建模能力、灵活声学特征提取能力及低延迟流式合成能力。其语音编码帧率为16.67 Hz,最低码率为0.43 kbps,最高码率为0.87 kbps。评估结果表明,LongCat-Audio-Codec在极低码率下仍保持良好语音可懂性,并能生成高质量语音,有效平衡编码效率与解码质量。推理代码与模型检查点已开源:https://github.com/meituan-longcat/LongCat-Audio-Codec。
原文摘要 · Abstract (English)
This paper presents LongCat-Audio-Codec, an audio tokenizer and detokenizer solution designed for industrial grade end-to-end speech large language models. By leveraging a decoupled model architecture and a multistage training strategy, LongCat-Audio-Codec exhibits robust semantic modeling capabilities, flexible acoustic feature extraction capabilities, and low-latency streaming synthesis capabilities. It encodes speech at an ultra-low frame rate of 16.67 Hz, with a minimum bitrate of 0.43 kbps and a maximum bitrate of 0.87 kbps. Evaluation results demonstrate that LongCat-Audio-Codec achieves strong speech intelligibility and is capable of synthesizing highquality speech at low bitrate, thus effectively balancing coding efficiency and decoding quality. The inference code and model checkpoints of LongCat-Audio-Codec are available at: https://github.com/meituan-longcat/LongCat-Audio-Codec.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。