arXiv:2606.09962cs.LGcs.AI2026-06

FSQ分词法让离散数据扩散模型更优,文本转语音效果更好且更快。

Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-Speech

论文配图:Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-Speech
图 1 · 摘自论文原文
  • 用KL散度和预测准确率分析潜空间结构,验证FSQ最优性。
  • 基于FSQ的语音合成模型性能超越主流LLM方案,体积更小速度更快。
  • 适合追求高效高质离散生成的开发者与研究者参考。

连续扩散模型用于离散数据生成是扩散模型家族的新方向,旨在探索自回归大语言模型的替代方案。本文通过严格的理论分析与数值实验,研究了基于Kullback-Leibler散度的扩散路径测度与最优训练模型对正确令牌预测的准确性,发现FSQ分词方案在潜空间结构上具有最适配连续扩散生成离散数据的特性。为验证实际效果,我们训练多个以语音标记为中间声学特征的文本到语音扩散模型,结果表明基于FSQ的模型表现最佳,且优于其强大的基于LLM的基线模型,同时模型体积更小、推理速度更快。

原文摘要 · Abstract (English)

Continuous diffusion for categorical data is a framework belonging to the diffusion family and aiming at generating discrete data. The scientific interest to such models has been constantly increasing these days because researchers try to achieve a challenging goal of finding reasonable alternatives to autoregressive large language models. In this paper, we study the properties of the structure of the latent space corresponding to discrete tokens expressed in terms of Kullback-Leibler divergence on diffusion path measures and accuracy of the correct token prediction by the optimally trained diffusion model. We find that FSQ tokenization scheme has the latent space structure with the properties that make it best suited for continuous diffusion for categorical data as verified through rigorous theoretical analysis and numerical experiments. To validate our findings in real-life scenario, we train several text-to-speech diffusion models having speech tokens as intermediate acoustic features, and show that the one based on FSQ tokens indeed performs the best, and, moreover, it outperforms its strong LLM-based counterpart, at the same time being significantly smaller and faster.

扩散模型语音合成离散生成FSQ

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。