arXiv:2508.05835eess.AScs.CL2025-08中稿 · Interspeech 2025被引 18

提出12.5帧/秒的超快语音编码器,实现高质量低延迟语音大模型推理。

NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference

  • 采用极低帧率设计,仅12.5帧/秒,大幅减少自回归生成步数。
  • 在多个码率下超越现有方法,在保持高音质的同时显著提升速度。
  • 适合追求低延迟语音生成与高效训练的语音大模型研究者使用。

大型语言模型(LLMs)通过音频编码器将语音离散化为标记,使语言建模技术可应用于语音数据。然而,现有音频编码器通常帧率较高,导致训练和推理速度慢,尤其对自回归模型影响显著。为此,学界日益关注低帧率音频编码器,以减少生成一秒钟音频所需的自回归步骤。本文通过消融实验研究帧率、码率与因果性对编码重建质量的影响。基于发现,我们提出NanoCodec,一种前沿音频编码器,可在仅12.5帧/秒的帧率下实现高质量压缩。NanoCodec在多种码率范围内均优于现有方法,为低延迟、高效的语音大模型训练与推理树立了新基准。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have significantly advanced audio processing by leveraging audio codecs to discretize audio into tokens, enabling the application of language modeling techniques to speech data. However, existing audio codecs often operate at high frame rates, leading to slow training and inference, particularly for autoregressive models. To address this, there is growing interest in low frame-rate audio codecs, which reduce the number of autoregressive steps required to generate one second of audio. In this paper, we conduct ablation studies to examine the impact of frame rate, bitrate, and causality on codec reconstruction quality. Based on our findings, we introduce NanoCodec, a state-of-the-art audio codec that achieves high-quality compression at just 12.5 frames per second (FPS). NanoCodec outperforms related works across various bitrate ranges, establishing a new benchmark for low-latency and efficient Speech LLM training and inference.

语音生成编码器大模型推理低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。