研究语音模型离散单元表示,发现不同模型规模需匹配不同聚类策略。
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
- 通过系统实验分析编码器与聚类粒度对预训练的影响。
- 揭示了离散词汇在语言与非语言特征中的有效使用模式。
- 强调数据聚类需与下游任务领域匹配以提升鲁棒性。
本文研究语音语言模型(SLMs)中离散单元表示的特性,重点优化持续预训练阶段的语音建模能力。我们系统考察了模型架构、数据表示及训练鲁棒性对预训练过程的影响,该过程旨在将现有预训练语言模型适配至语音模态。实验表明,语音编码器的选择和聚类粒度随模型规模变化而影响最优离散化策略。通过分析聚类分布与音素对齐情况,我们揭示了离散词汇在语言与副语言模式中的有效利用机制。此外,还探讨了聚类数据选择对模型鲁棒性的影响,强调离散化训练数据与目标应用领域的匹配至关重要。
原文摘要 · Abstract (English)
This paper investigates discrete unit representations in Speech Language Models (SLMs), focusing on optimizing speech modeling during continual pre-training. In this paper, we systematically examine how model architecture, data representation, and training robustness influence the pre-training stage in which we adapt existing pre-trained language models to the speech modality. Our experiments highlight the role of speech encoders and clustering granularity across different model scales, showing how optimal discretization strategies vary with model capacity. By examining cluster distribution and phonemic alignments, we investigate the effective use of discrete vocabulary, uncovering both linguistic and paralinguistic patterns. Additionally, we explore the impact of clustering data selection on model robustness, highlighting the importance of domain matching between discretization training and target applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。