小数据也能学出好音乐特征,挑战了大模型需海量数据的假设。
Learning Music Audio Representations With Limited Data
- 在5到8000分钟数据上测试多种模型,研究小样本学习表现
- 部分小数据训练模型性能接近大规模数据模型,随机初始化也有效
- 适合研究冷门音乐、个性化创作等数据稀缺场景
针对音乐音频表示学习,当前普遍认为需大量训练数据才能获得高性能。若属实,则在非主流音乐传统、小众流派或个性化音乐创作等数据稀缺场景中将面临挑战。本文系统考察多种音乐音频表示模型在小样本学习条件下的表现。实验涵盖不同架构、训练范式和输入时长的模型,在5至8,000分钟的数据集上进行训练,并在多个音乐信息检索任务上评估其性能,同时分析对噪声的鲁棒性。结果表明,在特定条件下,小数据训练的模型甚至随机初始化模型所学表示,性能可与大规模数据模型相当,但手工特征在部分任务上仍优于所有学习型表示。
原文摘要 · Abstract (English)
Large deep-learning models for music, including those focused on learning general-purpose music audio representations, are often assumed to require substantial training data to achieve high performance. If true, this would pose challenges in scenarios where audio data or annotations are scarce, such as for underrepresented music traditions, non-popular genres, and personalized music creation and listening. Understanding how these models behave in limited-data scenarios could be crucial for developing techniques to tackle them. In this work, we investigate the behavior of several music audio representation models under limited-data learning regimes. We consider music models with various architectures, training paradigms, and input durations, and train them on data collections ranging from 5 to 8,000 minutes long. We evaluate the learned representations on various music information retrieval tasks and analyze their robustness to noise. We show that, under certain conditions, representations from limited-data and even random models perform comparably to ones from large-dataset models, though handcrafted features outperform all learned representations in some tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。