基于公开编辑元数据构建大规模音乐版本识别数据集
Discogs-VI: A Musical Version Identification Dataset Based on Public Editorial Metadata
- 利用Discogs元数据构建超大版本数据集,覆盖近200万版本
- 通过精准搜索匹配YouTube官方上传,获得9.8万歌单、49.3万版本
- 无需复杂模型或数据增强,基线模型在多个数据集表现优异
当前版本识别(VI)数据集规模不足且音乐多样性有限,难以训练鲁棒神经网络;同时其非代表性的群组大小分布导致系统评估不真实。为解决这些问题,本文探索了Discogs音乐数据库中丰富的编辑元数据潜力,构建了一个包含约190万版本、跨越34.8万个群组的大规模版本数据集。通过高精度搜索算法,将该数据集映射至YouTube官方音乐上传内容,最终形成约49.3万版本、9.8万个群组的数据集。相比现有数据集,本数据集的群组数量增加超过9倍,版本数量增加超过4倍。我们通过训练一个基础神经网络(无需复杂结构或数据增强),在SHS100K和Da-TACOS数据集上取得了具有竞争力的结果。所构建数据集、工具、提取的音频特征及训练模型均已公开发布。
原文摘要 · Abstract (English)
Current version identification (VI) datasets often lack sufficient size and musical diversity to train robust neural networks (NNs). Additionally, their non-representative clique size distributions prevent realistic system evaluations. To address these challenges, we explore the untapped potential of the rich editorial metadata in the Discogs music database and create a large dataset of musical versions containing about 1,900,000 versions across 348,000 cliques. Utilizing a high-precision search algorithm, we map this dataset to official music uploads on YouTube, resulting in a dataset of approximately 493,000 versions across 98,000 cliques. This dataset offers over nine times the number of cliques and over four times the number of versions than existing datasets. We demonstrate the utility of our dataset by training a baseline NN without extensive model complexities or data augmentations, which achieves competitive results on the SHS100K and Da-TACOS datasets. Our dataset, along with the tools used for its creation, the extracted audio features, and a trained model, are all publicly available online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。