用无标签音频自动构建鼓音色库,生成高质量合成数据训练打击乐转录模型。
Towards Realistic Synthetic Data for Automatic Drum Transcription
- 从无标签音频中半监督提取多样鼓音色样本,构建大规模音色库。
- 仅用MIDI文件合成高质量音频数据,模型在两个基准上达新最优。
- 适合缺乏配对数据的音乐信息检索与自动转录研究者使用。
深度学习模型在自动打击乐转录(ADT)中表现领先,但其性能依赖于大规模配对音轨-MIDI数据集,这类数据稀缺。现有合成数据方法常因使用低保真SoundFont库而引入显著领域差距。尽管高质量单次采样更具优势,却缺乏标准化、大规模的可用格式。本文提出一种新范式,无需配对音频-MIDI训练数据。核心贡献是提出一种半监督方法,从无标签音频中自动收集大规模且多样的单次鼓音色样本。利用该音色库,仅凭MIDI文件即可合成高质量数据集,并用于训练序列到序列的转录模型。我们在ENST和MDB测试集上评估,结果显著优于全监督方法及先前合成数据方法,达到新最优水平。实验代码已公开于https://github.com/pier-maker92/ADT_STR。
原文摘要 · Abstract (English)
Deep learning models define the state-of-the-art in Automatic Drum Transcription (ADT), yet their performance is contingent upon large-scale, paired audio-MIDI datasets, which are scarce. Existing workarounds that use synthetic data often introduce a significant domain gap, as they typically rely on low-fidelity SoundFont libraries that lack acoustic diversity. While high-quality one-shot samples offer a better alternative, they are not available in a standardized, large-scale format suitable for training. This paper introduces a new paradigm for ADT that circumvents the need for paired audio-MIDI training data. Our primary contribution is a semi-supervised method to automatically curate a large and diverse corpus of one-shot drum samples from unlabeled audio sources. We then use this corpus to synthesize a high-quality dataset from MIDI files alone, which we use to train a sequence-to-sequence transcription model. We evaluate our model on the ENST and MDB test sets, where it achieves new state-of-the-art results, significantly outperforming both fully supervised methods and previous synthetic-data approaches. The code for reproducing our experiments is publicly available at https://github.com/pier-maker92/ADT_STR
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。