arXiv:2606.23285cs.CLcs.SD2026-06中稿 · Interspeech2026

降低语音离散表示的比特率,仍能生成自然语音,且续写质量稳定。

On the Effect of Segmentation Width and Cluster Size on Speech Resynthesis and Continuation in Generative Spoken Language Models

  • 用固定宽度分段+多聚类数训练,实现多种比特率设置下的语音建模。
  • 在低于基线的比特率下仍可生成可懂且自然的语音,续写质量保持稳定。
  • 提示现有自动评估方法不稳定,适合关注语音生成效率的研究者。

生成式口语语言模型(GSLM)通过使用离散语音表示而非文本转录来实现无文本语音建模。本文研究了在不同比特率下,基于离散语音表示的语音合成与续写性能。采用固定宽度分段,并在多个聚类数下训练K-means模型,生成多种比特率配置。结果表明,可在低于基线的比特率下生成可懂且自然的语音。此外,在多个指标上,低比特率下的语音续写质量依然稳定,说明传统GSLM设置可能对有效语音生成而言冗余。尽管基于大语言模型的指标比传统指标与人类主观评分相关性更高,但相关性仍较低,凸显了更稳定自动评估方法的必要性。

原文摘要 · Abstract (English)

Generative Spoken Language Modeling (GSLM) enables text-free speech modeling by training language models (LMs) using discrete speech representations instead of textual transcription. In this paper, we investigate the performance of GSLM on speech synthesis and continuation using discrete speech representations with varying bitrates. We segment speech representations with fixed widths and train K-means models in multiple cluster sizes, resulting in various bitrate settings. We demonstrate that intelligible and natural speech can be synthesized at lower bitrate settings than the baseline. Furthermore, speech continuation quality remains stable at lower bitrates across multiple metrics, suggesting that the conventional GSLM setting may be redundant for effective speech generation. Although LLM-based metrics show higher correlation with human subjective score than conventional metrics, it remains low, highlighting the need for more stable automatic evaluation methods.

语音生成离散表示比特率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。