arXiv:2604.24416cs.CLcs.AI2026-04

连续扩散模型让语音语言模型更高效,可生成多语种多人声情感语音。

Scaling Properties of Continuous Diffusion Spoken Language Models

论文配图:Scaling Properties of Continuous Diffusion Spoken Language Models
图 1 · 摘自论文原文
  • 用连续扩散框架替代离散自回归,减少计算瓶颈。
  • 160亿参数模型在数千万小时数据下生成带语调的多说话人语音。
  • 计算量增大时模型性能趋于稳定,适合快速推理。

仅基于语音的语音语言模型(SLMs)性能落后于文本及文-音模型,近期离散自回归(AR)SLMs表明其需巨大算力与数据才能接近文本模型水平。由于离散化语音会带来瓶颈,本文探索连续扩散(CD)SLM的可行性。为量化语言质量,提出音素杰恩-申农散度(pJSD)指标。分析显示,CD SLMs呈现与AR模型相似的缩放规律:验证损失和pJSD随规模增长而下降,且最优参数-令牌比随算力增加而降低。当算力提升后,损失对数据和模型大小不敏感,具备快速推理潜力。将CD SLM扩展至160亿参数,使用数千万小时对话数据,可生成富有情感、语调自然、多说话人、多语言的语音,但长篇连贯性仍是重大挑战。

原文摘要 · Abstract (English)

Speech-only spoken language models (SLMs) lag behind text and text-speech models in performance, with recent discrete autoregressive (AR) SLMs indicating significant computational and data demands to match text models. Since discretizing continuous speech for AR creates bottlenecks, we explore whether continuous diffusion (CD) SLM is more viable. To quantify the SLMs linguistic quality, we introduce the phoneme Jensen-Shannon divergence (pJSD) metric. Our analysis reveals CD SLMs, mirroring AR behavior, exhibit scaling laws for validation loss and pJSD, and show optimal token-to-parameter ratios decreasing as compute scales. However, for the latter, loss becomes insensitive to choice of data and model sizes, showing potential for fast inference. Scaling CD SLMs to 16B parameters with tens of millions of hours of conversational data enables generation of emotive, prosodic, multi-speaker, multilingual speech, though achieving long-form coherence remains a significant challenge.

语音生成扩散模型语言模型连续建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。