arXiv:2602.06602cs.SDcs.AI2026-02被引 5

用扩散模型实现低比特率语音分词,兼顾理解与还原质量。

Scaling Speech Tokenizers with Diffusion Autoencoders

  • 基于扩散自编码器联合学习语义与声学特征。
  • 在200 bit/s下实现12.5 Hz的极低令牌率。
  • 适合语音生成、压缩及低带宽场景应用。

语音分词是语音语言模型的基础,但现有方法面临两大挑战:(1) 在理解所需的语义编码与重建所需的声学编码之间难以平衡;(2) 难以同时实现低比特率和低令牌率。我们提出语音扩散分词器(SiTok),一种扩散自编码器,通过监督学习联合学习语义丰富的表示,并利用扩散模型实现高保真音频重建。我们将SiTok扩展至16亿参数,并在200万小时语音数据上进行训练。实验表明,SiTok在理解、重建和生成任务上均优于强基线,在12.5赫兹的极低令牌率和200比特每秒的比特率下表现卓越。

原文摘要 · Abstract (English)

Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understanding and acoustics for reconstruction, and (2) achieving low bit rates and low token rates. We propose Speech Diffusion Tokenizer (SiTok), a diffusion autoencoder that jointly learns semantic-rich representations through supervised learning and enables high-fidelity audio reconstruction with diffusion. We scale SiTok to 1.6B parameters and train it on 2 million hours of speech. Experiments show that SiTok outperforms strong baselines on understanding, reconstruction and generation tasks, at an extremely low token rate of $12.5$ Hz and a bit-rate of 200 bits-per-second.

语音分词扩散模型低码率自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。