arXiv:2601.09239cs.SDcs.AI2026-01被引 2

将语音拆分为语义与音色独立的离散编码,提升语音生成质量与可控性。

DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion

  • 通过不同优化目标分离语义与音色信息,实现更彻底的解耦。
  • 支持高保真重建和跨说话人语音克隆,语音识别错误率低。
  • 适合语音合成、语音克隆等需精确控制音色的下游任务。

语音分词器是全离散语音大模型的关键组件。现有分词器或侧重语义编码,或将语义与音色混合,或仅实现不完全的语义-音色解耦。为实现更好解耦,本文提出DSA-Tokenizer,通过不同优化约束显式将语音分解为离散的语义与音色令牌。其中,语义令牌由语音识别(ASR)监督以捕捉语言内容,音色令牌则专注于梅尔频谱图重建以编码发音风格。我们进一步引入层次化流匹配解码器与联合重建-上下文填空训练策略,使模型同时支持高保真重建与跨语音段语音克隆。为加速推理,我们将解码器蒸馏至4步生成,并通过GAN微调提升合成质量。实验表明,DSA-Tokenizer在语义-音色解耦、可控语音克隆与高效高保真生成方面表现优异,且具有低词错误率(WER)与字符错误率(CER)。结果还表明,解耦分词为下游大模型语音生成提供了更有效的接口。音频样例见:https://anonymous.4open.science/w/DSA_Tokenizer_demo/

原文摘要 · Abstract (English)

Speech tokenizers are a key building block of fully discrete Speech LLMs. Existing tokenizers either prioritize semantic encoding, fuse semantic content with acoustic style inseparably, or achieve incomplete semantic-acoustic disentanglement. To achieve better disentanglement, we propose DSA-Tokenizer, which explicitly disentangles speech into discrete semantic and acoustic tokens via distinct optimization constraints. Specifically, semantic tokens are supervised by ASR to capture linguistic content, while acoustic tokens focus on mel-spectrograms restoration to encode style. We further introduce a hierarchical Flow Matching decoder and a joint reconstruction-context inpainting training strategy, allowing the model to support both high-fidelity reconstruction and cross-utterance voice clone. To speed up inference, we distill the dit decoder to 4-step inference and improve synthesis quality with GAN fine-tuning. Experiments demonstrate that DSA-Tokenizer provides strong semantic-acoustic disentanglement, reliable controllable voice cloning, and efficient high-fidelity generation with low WER/CER. Moreover, our results suggest that disentangled tokenization provides a more effective interface for downstream large-model speech generation. Audio samples are avaialble at https://anonymous.4open.science/w/DSA_Tokenizer_demo/

语音生成解耦表征语音克隆离散编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。