arXiv:2511.10232cs.CLcs.AI2025-11

用多码本+多标记预测,让语音模型响应快一倍

VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction

  • 用多码本编码直接生成语音标记,跳过耗时的合成步骤
  • 首块延迟从725毫秒降至350毫秒,性能不降
  • 适合需要实时交互的语音应用开发

当前端到端语音语言模型虽有进展,但仍存在显著响应延迟,主要源于语音标记的自回归生成和复杂流匹配模型的依赖。为解决此问题,我们提出VocalNet-M2,一种融合多码本分词与多标记预测(MTP)策略的低延迟语音模型。该模型直接生成多码本语音标记,避免了高延迟的流匹配模型。MTP策略提升了生成效率并改善整体性能。大量实验表明,VocalNet-M2将首块延迟从约725毫秒降至350毫秒,同时在主流语音语言模型上保持竞争力。本工作还系统对比了单码本与多码本策略,为实时交互应用中的高效高性能语音模型开发提供了重要参考。

原文摘要 · Abstract (English)

Current end-to-end spoken language models (SLMs) have made notable progress, yet they still encounter considerable response latency. This delay primarily arises from the autoregressive generation of speech tokens and the reliance on complex flow-matching models for speech synthesis. To overcome this, we introduce VocalNet-M2, a novel low-latency SLM that integrates a multi-codebook tokenizer and a multi-token prediction (MTP) strategy. Our model directly generates multi-codebook speech tokens, thus eliminating the need for a latency-inducing flow-matching model. Furthermore, our MTP strategy enhances generation efficiency and improves overall performance. Extensive experiments demonstrate that VocalNet-M2 achieves a substantial reduction in first chunk latency (from approximately 725ms to 350ms) while maintaining competitive performance across mainstream SLMs. This work also provides a comprehensive comparison of single-codebook and multi-codebook strategies, offering valuable insights for developing efficient and high-performance SLMs for real-time interactive applications.

语音生成低延迟多码本序列预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。