选更长的语音片段,能用一半数据达到更好效果。
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
- 优先选取最长语音片段,而非追求多样性或总量。
- 仅用原数据一半,ASR性能反而提升,训练时间减少24%。
- 适合想高效训练语音模型的研究者与工程师。
自监督学习(SSL)已重塑语音处理,但其对大规模预训练数据集的依赖仍是瓶颈。尽管鲁棒性常归因于数据规模与多样性,但数据分布的作用仍不明确。本文系统研究了预训练数据子集对自动语音识别(ASR)性能的影响。令人惊讶的是,针对声学、说话人或语言多样性优化的数据选择并未带来明显提升,反而随机采样表现相当。相反,优先选取最长语音片段可显著提升ASR性能,且仅需使用原数据的一半,使大型语料库上的预训练时间减少24%。结果表明,在预训练语音SSL模型时,数据长度比多样性或总量更为关键,为SSL语音处理中的数据选择策略提供了新视角。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) has transformed speech processing, yet its reliance on massive pre-training datasets remains a bottleneck. While robustness is often attributed to scale and diversity, the role of the data distribution is less understood. We systematically examine how curated subsets of pre-training data influence Automatic Speech Recognition (ASR) performance. Surprisingly, optimizing for acoustic, speaker, or linguistic diversity yields no clear improvements over random sampling. Instead, we find that prioritizing the longest utterances achieves superior ASR results while using only half the original dataset, reducing pre-training time by 24% on a large corpora. These findings suggest that for pre-training speech SSL models, data length is a more critical factor than either data diversity or overall data quantity for performance and efficiency, offering a new perspective for data selection strategies in SSL speech processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。