低帧率语音编解码器提升大模型语音训练与推理速度
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
- 用有限标量量化和对抗训练实现低帧率编码
- 1.89 kbps比特率下保持21.5帧/秒,音质接近现有模型
- 使语音大模型推理提速3倍,适合高效语音生成场景
大型语言模型(LLMs)通过将音频转换为离散标记的音频编解码器,在音频处理中取得显著进展,使语言建模技术可应用于音频数据。然而,传统音频编解码器通常帧率较高,导致自回归模型训练与推理速度缓慢。为此,我们提出低帧率语音编解码器(LFSC):一种基于有限标量量化和与大型语音语言模型联合对抗训练的神经音频编解码器,可在1.89 kbps比特率下实现21.5帧/秒的编码速率,同时保持高质量音频压缩。实验表明,该编解码器可使基于LLM的文本到语音模型推理速度提升约3倍,同时改善语音可懂度,且音质与先前模型相当。
原文摘要 · Abstract (English)
Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and inference, especially for autoregressive models. To address this challenge, we present the Low Frame-rate Speech Codec (LFSC): a neural audio codec that leverages finite scalar quantization and adversarial training with large speech language models to achieve high-quality audio compression with a 1.89 kbps bitrate and 21.5 frames per second. We demonstrate that our novel codec can make the inference of LLM-based text-to-speech models around three times faster while improving intelligibility and producing quality comparable to previous models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。