自监督音频模型可灵活学习跨领域声学特征,无需标签即可迁移使用。
Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study
- 用BYOL-A方法在语音/非语音数据上预训练卷积模型
- 不同预训练数据的模型在多数下游任务中表现接近甚至超越专用模型
- 适合缺乏标注数据的跨域迁移或探索性研究
自监督学习(SSL)算法能利用大量无标签音频数据预训练出鲁棒的表示,支持多种下游任务。以往这些方法多针对语音或非语音场景分别开发。本文采用自监督预训练方法(BYOL-A),研究了卷积模型在不同下游任务中对预训练数据领域的敏感性。结果表明,无论在语音、非语音数据,或两者混合数据上预训练的模型,均能在几乎全部下游任务中取得优异表现,性能优于或接近主流领域专用模型。不同预训练数据间的领域特异性差异很小。尽管领域专用模型在其目标领域表现极佳,但在其他领域普遍下降。这些结果说明,SSL方法可有效学习通用性强的声学表示,无需标签即可用于后续迁移学习、微调或数据探索,尤其适用于下游数据与预训练数据相似或存在领域差异的情况。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) algorithms have emerged as powerful tools that can leverage large quantities of unlabeled audio data to pre-train robust representations that support strong performance on diverse downstream tasks. Up to now these have mostly been developed separately for speech and non-speech applications. Here, we explored the domain specificity of a convolutional model's pre-training data relative to different downstream speech and non-speech tasks using a self-supervised pre-training approach (BYOL-A). We found that these pre-trained models (regardless of whether they were pre-trained on speech data, non-speech data or both) enabled good performance on nearly all downstream tasks, beating or nearly matching the performance of popular domain-specific models. Only small domain-specificity advantages were observed between the different pre-training datasets. The popular domain-specific models used as baselines performed very well in their target domains, but generally faltered outside of them. Together, these results demonstrate that SSL methods can be a powerful way to learn flexible representations for domain specific data without labels. These models can be a powerful resource for later transfer learning, fine-tuning or data exploration applications when the downstream data are similar, but also perhaps when there may be a domain mismatch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。