arXiv:2409.10831cs.SDcs.AI2024-09中稿 · 2025 IEEE Internat…被引 21

构建了25万首免费可商用的乐谱数据集,助力音乐AI公平发展。

PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing

  • 从MuseScore收集25万首公共领域乐谱,全部免费可商用。
  • 通过用户评分等元数据筛选高质量乐谱,提升数据可信度。
  • 适合研究音乐生成、数据质量评估的AI开发者与学者使用。

近年来生成式音乐AI的爆发引发了数据版权、音乐人授权及开源AI与大型企业之间的矛盾。这些问题凸显了公开可用、无版权音乐数据的迫切需求,尤其在符号化音乐数据方面尤为匮乏。为缓解此问题,我们提出了PDMX:一个来自MuseScore乐谱分享论坛的超大规模开源音乐数据集,包含超过25万首公共领域MusicXML乐谱,据我们所知是目前最大的免费符号音乐数据集。PDMX还包含丰富的标签和用户互动元数据,可有效分析并筛选高质量的用户生成乐谱。基于该数据收集流程提供的附加元数据,我们开展了多轨音乐生成实验,评估不同子集对下游模型行为的影响,并验证用户评分统计可作为数据质量的有效衡量指标。示例见 https://pnlong.github.io/PDMX.demo/。

原文摘要 · Abstract (English)

The recent explosion of generative AI-Music systems has raised numerous concerns over data copyright, licensing music from musicians, and the conflict between open-source AI and large prestige companies. Such issues highlight the need for publicly available, copyright-free musical data, in which there is a large shortage, particularly for symbolic music data. To alleviate this issue, we present PDMX: a large-scale open-source dataset of over 250K public domain MusicXML scores collected from the score-sharing forum MuseScore, making it the largest available copyright-free symbolic music dataset to our knowledge. PDMX additionally includes a wealth of both tag and user interaction metadata, allowing us to efficiently analyze the dataset and filter for high quality user-generated scores. Given the additional metadata afforded by our data collection process, we conduct multitrack music generation experiments evaluating how different representative subsets of PDMX lead to different behaviors in downstream models, and how user-rating statistics can be used as an effective measure of data quality. Examples can be found at https://pnlong.github.io/PDMX.demo/.

音乐数据集符号音乐开源数据AI生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。