arXiv:2505.17446cs.CLcs.SD2025-05中稿 · Interspeech2025被引 6

探索分割粒度与词表大小对语音建模的影响,发现适度粗粒度更高效。

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

  • 采用固定/可变分割和多簇K-means聚类生成离散语音表示
  • 中等粗粒度分割与更大词表使零样本语音理解性能提升,训练效率提高70%
  • 适合关注语音令牌化设计与训练加速的研究者

语音分词旨在将语音信号转化为离散表示序列,是语音语言模型(SLMs)的基础。本文研究了分词的两个关键因素:分割宽度与离散单元的聚类数量。首先,将语音信号按固定或可变宽度进行分段,并生成聚合表示;随后在多种聚类数量下训练K-means模型。在零样本语音理解基准上的评估显示,适度粗粒度分割与更大的聚类数量具有正向效果。值得注意的是,表现最佳的模型中,最高效的方案实现了50%的训练数据减少和70%的训练时间降低。分析表明,组合多个标记有助于提升细粒度语音理解能力。

原文摘要 · Abstract (English)

The purpose of speech tokenization is to transform a speech signal into a sequence of discrete representations, serving as the foundation for speech language models (SLMs). While speech tokenization has many options, their effect on the performance of SLMs remains unclear. This paper investigates two key aspects of speech tokenization: the segmentation width and the cluster size of discrete units. First, we segment speech signals into fixed/variable widths and pooled representations. We then train K-means models in multiple cluster sizes. Through the evaluation on zero-shot spoken language understanding benchmarks, we find the positive effect of moderately coarse segmentation and bigger cluster size. Notably, among the best-performing models, the most efficient one achieves a 50% reduction in training data and a 70% decrease in training runtime. Our analysis highlights the importance of combining multiple tokens to enhance fine-grained spoken language understanding.

语音建模分词优化训练效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。