arXiv:2607.08756cs.SDcs.LG2026-07

构建了572段流行音乐多轨乐谱数据集,用于评估自动编曲转录模型。

MulTTiPop: A Multitrack Transcription Dataset for Pop Music

论文配图:MulTTiPop: A Multitrack Transcription Dataset for Pop Music
图 1 · 摘自论文原文
  • 基于元数据匹配音频与MIDI,人工对齐节拍后进行时间对齐
  • 包含3.5小时音乐,覆盖1930至2000年代多种流派,挑战现有模型
  • 最佳模型仅达29%音符起始检测准确率,表明领域仍需突破

我们提出MulTTiPop,一个用于评估自动音乐转录模型的流行音乐片段及其多轨MIDI录音数据集。该数据集包含572个音乐片段,总时长3.5小时,涵盖从1930年代到2000年代的多种音乐风格。通过在Lakh MIDI和TheoryTab数据集中基于元数据匹配歌曲片段,人工识别音频与MIDI之间的锚定节拍,再对音频进行节拍跟踪,并将MIDI进行时间扭曲以匹配其节奏与时间。我们在MulTTiPop上评估了当前最先进的自动音乐转录模型,发现仍有巨大提升空间,最佳模型仅达到29%的起始点F1值。更多细节与音频示例可访问https://gclef-cmu.org/multtipop。

原文摘要 · Abstract (English)

We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. To collect this dataset, we perform metadata-based matching on song segments from the Lakh MIDI and TheoryTab datasets, manually identify an anchor beat between the audio and MIDI, then use beat tracking on the audio and warp the MIDI to match its tempo and timing. We evaluate state-of-the-art automatic music transcription models on MulTTiPop and find substantial room for improvement, with the best model achieving 29% Onset F1. More details and sound examples of MulTTiPop are available at https://gclef-cmu.org/multtipop.

音乐转录多轨数据集流行音乐自动作曲

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。