arXiv:2506.22362eess.AScs.LG2025-06

用扩散模型提升语音分词效率,50词/秒质量媲美双倍速率标准模型

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

  • 用语义词条件化神经编码器,减少语义与声学词冗余
  • 50词/秒时语音质量达标准SoundStream两倍速率水平
  • 仅4步扩散采样即可完成高质量波形生成,适合高效语音合成

基于词的语音建模是语音生成的重要方法,通过自监督学习(SSL)模型提取特征并经神经语音编码器量化得到语义词和声学词。这些词通常采用自回归方式建模,推理速度受制于词率。本文提出DiffSoundStream,在非流式场景下通过两项技术提升语音分词效率:(1) 将神经编码器以语义词为条件,降低语义与声学词间的冗余;(2) 利用潜在扩散模型从语义词和粗粒度声学词中合成高质量波形。实验表明,在50词/秒时,DiffSoundStream的语音质量与运行在两倍词率下的标准SoundStream模型相当。此外,仅需4步扩散采样即可实现步长蒸馏,质量损失极小。

原文摘要 · Abstract (English)

Token-based language modeling is a prominent approach for speech generation, where tokens are obtained by quantizing features from self-supervised learning (SSL) models and extracting codes from neural speech codecs, generally referred to as semantic tokens and acoustic tokens. These tokens are often modeled autoregressively, with the inference speed being constrained by the token rate. In this work, we propose DiffSoundStream, a solution that improves the efficiency of speech tokenization in non-streaming scenarios through two techniques: (1) conditioning the neural codec on semantic tokens to minimize redundancy between semantic and acoustic tokens, and (2) leveraging latent diffusion models to synthesize high-quality waveforms from semantic and coarse-level acoustic tokens. Experiments show that at 50 tokens per second, DiffSoundStream achieves speech quality on par with a standard SoundStream model operating at twice the token rate. Additionally, we achieve step-size distillation using just four diffusion sampling steps with only a minor quality loss.

语音生成扩散模型分词效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。