HoliTok将语音转为高效隐向量,支持生成与理解统一建模。
HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

- 用连续隐变量编码48kHz语音,输出25Hz、128维的紧凑序列
- 在统一架构中实现高质量语音合成与识别,无需额外优化技巧
- 首次在生成与理解任务中同时保持高保真与强可学习性
统一语音基础模型需要一个既可被语言模型学习又可解码为高质量波形的整体化分词空间。现有语音分词器常难以同时满足这两个要求,导致结构复杂且训练设计繁琐。我们提出HoliTok,一种面向统一生成-理解建模的连续整体语音分词模型。该模型将48~kHz语音编码为25~Hz、128维的紧凑隐向量序列,通过渐进式训练策略联合保持信号级保真度、融入语义信息并维持强隐向量可学习性。基于此分词结果,构建了统一的AR+DiT模型,用于语音合成与识别,同一隐向量序列同时支持生成专用与统一生成-理解任务。实验表明,HoliTok实现了有竞争力的重建保真度,提升了高质量可控合成的生成可学习性,并在所评估表示中唯一能在统一生成-理解架构中稳健运行而无需额外优化技巧。结果表明,HoliTok可作为有效的语音分词器及统一语音建模的基础表征接口。代码已公开:https://github.com/bovod-sjtu/HoliTok。
原文摘要 · Abstract (English)
Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenizers, however, often fail to satisfy these requirements simultaneously, leading to increased architectural complexity and more involved training designs. We propose HoliTok, a continuous Holistic speech Tokenization model designed for unified generation-understanding modeling. HoliTok encodes 48~kHz speech into a compact 25~Hz sequence of 128-dimensional latents. It is trained with a progressive strategy that jointly preserves signal-level fidelity, incorporates semantic information, and maintains strong latent learnability. Based on this tokenization, we build a unified AR+DiT model for speech synthesis and recognition, where the same latent sequence supports both generation-specific and unified generation-understanding tasks. Experiments show that HoliTok achieves competitive reconstruction fidelity, improves generative learnability for high-quality and controllable synthesis, and, among the evaluated representations, is the only one that operates robustly in our unified generation-understanding architecture without additional optimization tricks. These results suggest that HoliTok serves as an effective speech tokenizer and a foundational representation interface for unified spoken language modeling. The code is available at: https://github.com/bovod-sjtu/HoliTok.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。