arXiv:2506.12537cs.CLcs.AI2025-06AAAI

优化语音分词器设计,提升大模型语音生成质量与速度

What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study

  • 采用解耦分词策略,增强语音与文本对齐
  • 多标记预测使推理速度提升12倍,字错误率降至3.01%
  • 引入角色感知生成与大规模角色问答数据集

语音-语言模型(SLMs)为统一语音与文本的理解和生成提供了前景。然而,跨模态对齐与高质量语音生成仍面临挑战。本文系统研究了大模型中心型SLMs中语音分词器设计的作用,结合语音头与说话人建模,在公平的SLM框架下对比了耦合、半解耦与全解耦分词器,发现解耦分词显著提升对齐与合成质量。为缓解语音与文本信息密度差异,提出多标记预测(MTP),使每个隐状态可解码多个语音标记,实现最高12倍的解码加速,并将词错误率从6.07降至3.01。此外,提出说话人感知生成范式,构建了包含多样说话人身份的大规模角色扮演知识问答基准RoleTriviaQA。实验表明,该方法同时提升了知识理解与说话人一致性。

原文摘要 · Abstract (English)

Speech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12$\times$ faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency.

语音生成大模型分词器多标记预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。