用聚类质量与嵌入秩预测语音自监督模型性能,1小时音频即可省下大量算力。
Towards Early Prediction of Self-Supervised Speech Model Performance
- 通过嵌入的聚类质量和秩评估预训练质量,无需依赖损失值。
- 仅用1小时无标签音频,相关性优于传统损失指标。
- 适合需要高效评估语音模型的科研与工程人员。
在自监督学习(SSL)中,预训练与评估耗费大量资源。语音领域当前的预训练质量指标(如损失值)与下游性能相关性差,难以在预训练阶段低成本预测最终表现。本文提出无监督高效方法,通过测量SSL语音模型嵌入的聚类质量与秩来评估预训练效果。实验表明,这些指标比预训练损失与下游性能的相关性更高,仅需1小时无标签音频即可实现有效评估,显著减少对GPU时长和标注数据的需求。
原文摘要 · Abstract (English)
In Self-Supervised Learning (SSL), pre-training and evaluation are resource intensive. In the speech domain, current indicators of the quality of SSL models during pre-training, such as the loss, do not correlate well with downstream performance. Consequently, it is often difficult to gauge the final downstream performance in a cost efficient manner during pre-training. In this work, we propose unsupervised efficient methods that give insights into the quality of the pre-training of SSL speech models, namely, measuring the cluster quality and rank of the embeddings of the SSL model. Results show that measures of cluster quality and rank correlate better with downstream performance than the pre-training loss with only one hour of unlabeled audio, reducing the need for GPU hours and labeled data in SSL model evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。