在算力有限下,优化语音大模型训练的架构与数据策略。
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
- 采用更轻量的模型结构,在相同算力下表现更优。
- 预训练数据量至关重要,数据迭代不足会显著降低性能。
- 发现模型大小与数据量的权衡关系,适合特定算力预算。
尽管基础模型取得了显著成功,但其训练仍需大量计算资源。本文研究在有限算力预算下,如何高效训练自监督语音基础模型。我们考察了影响算力消耗的关键因素,包括模型架构、模型规模和数据规模。通过在完全可比的设置下基准测试自监督学习目标,发现其他因素对自监督学习的成功影响更大。结果表明,在相同算力与参数预算下,更轻量的模型架构优于常见的小型架构。我们证明,即使在自监督训练中使用数据增强,预训练数据量依然关键,数据迭代次数过少会导致性能下降。最后,我们识别出模型大小与数据大小之间的权衡关系,揭示了给定算力预算下的最优模型规模。
原文摘要 · Abstract (English)
Despite their impressive success, training foundation models remains computationally costly. This paper investigates how to efficiently train speech foundation models with self-supervised learning (SSL) under a limited compute budget. We examine critical factors in SSL that impact the budget, including model architecture, model size, and data size. Our goal is to make analytical steps toward understanding the training dynamics of speech foundation models. We benchmark SSL objectives in an entirely comparable setting and find that other factors contribute more significantly to the success of SSL. Our results show that slimmer model architectures outperform common small architectures under the same compute and parameter budget. We demonstrate that the size of the pre-training data remains crucial, even with data augmentation during SSL training, as performance suffers when iterating over limited data. Finally, we identify a trade-off between model size and data size, highlighting an optimal model size for a given compute budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。