构建百万级钢琴乐谱数据集,用于符号音乐建模研究。
Aria-MIDI: A Dataset of Piano MIDI Files for Symbolic Music Modeling
- 用语言模型自动爬取网络音频并转录为MIDI文件
- 生成超百万个唯一MIDI文件,涵盖约十万小时音频
- 提供元数据标签和完整数据集,适合音乐生成研究
我们引入了一个大规模的MIDI文件数据集,通过将钢琴演奏音频录制转录为音符构成。该数据管道采用多阶段流程:首先利用语言模型基于元数据自主爬取互联网上的音频记录并评分,随后通过音频分类器进行筛选与分割。最终数据集包含超过一百万个独特MIDI文件,对应约十万小时的转录音频。我们对技术方法进行了深入分析,提供了统计洞察,并提取了元数据标签,一并开放。数据集可在 https://github.com/loubbrad/aria-midi 获取。
原文摘要 · Abstract (English)
We introduce an extensive new dataset of MIDI files, created by transcribing audio recordings of piano performances into their constituent notes. The data pipeline we use is multi-stage, employing a language model to autonomously crawl and score audio recordings from the internet based on their metadata, followed by a stage of pruning and segmentation using an audio classifier. The resulting dataset contains over one million distinct MIDI files, comprising roughly 100,000 hours of transcribed audio. We provide an in-depth analysis of our techniques, offering statistical insights, and investigate the content by extracting metadata tags, which we also provide. Dataset available at https://github.com/loubbrad/aria-midi.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。