arXiv:2511.13732eess.AScs.LG2025-11中稿 · ICASSP 2026被引 3

通过声学相似分组提升语音生成速度,保持音质不变。

Principled Coarse-Grained Acceptance for Speculative Decoding in Speech

  • 将语音令牌按声学相似性聚类,以组为单位验证草稿
  • 在LibriTTS上接受率和吞吐量显著提升,音质与说话人特征保持稳定
  • 适合需要高速语音生成且对质量要求高的场景

推测解码通过快速草稿模型提出候选词,由大目标模型验证来加速自回归语音生成。然而,对于生成声学标记的语音大模型,严格的逐标记匹配过于苛刻:许多离散标记在声学或语义上可互换,导致接受率下降并限制加速效果。本文提出原理性粗粒度(PCG),基于目标模型嵌入空间构建声学相似性组(ASGs),将每个标记的概率质量分配到包含它的重叠组中,定义一种考虑重叠关系的粗粒度分布,并对所得组变量进行拒绝采样。该方法在组层面保证精确性,同时允许被接受的草稿标记代表组内任意成员。在LibriTTS数据集上,相较于标准推测解码及先前语音专用松弛方法,PCG提升了接受率和吞吐量,同时维持了语音可懂度与说话人相似性。结果表明,基于声学感知的组级接受是一种简单且通用的加速语音标记生成方式。

原文摘要 · Abstract (English)

Speculative decoding accelerates autoregressive speech generation by letting a fast draft model propose tokens that a larger target model verifies. However, for speech LLMs that generate acoustic tokens, exact token matching is overly restrictive: many discrete tokens are acoustically or semantically interchangeable, reducing acceptance rates and limiting speedups. We introduce Principled Coarse-Graining (PCG), which verifies proposals at the level of Acoustic Similarity Groups (ASGs) derived from the target model's embedding space. By splitting each token's probability mass across the overlapping groups that contain it, we define an overlap-aware coarse-grained distribution and perform rejection sampling on the resulting group variable. This yields an exactness guarantee at the group level while allowing the accepted draft token to stand in for any member of the group in practice. On LibriTTS, PCG increases acceptance and throughput relative to standard speculative decoding and prior speech-specific relaxations while maintaining intelligibility and speaker similarity. These results suggest acoustically aware, group-level acceptance as a simple and general way to accelerate speech token generation while maintaining speech quality.

语音生成推测解码声学相似性加速生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。