arXiv:2509.16662cs.SDcs.AI2025-09中稿 · publication at ISM…被引 2

清理了17万份MIDI数据中的3.8万份重复文件,提升训练可靠性。

On the de-duplication of the Lakh MIDI dataset

  • 用规则+对比学习方法识别重复音乐片段
  • 在LMD数据集中过滤出至少38,134个重复样本
  • 适合音乐信息检索与数据质量研究者

大规模数据集对训练泛化性强的深度学习模型至关重要。多数数据通过网络爬取获得,不可避免引入重复内容。在乐谱音乐领域,重复常源于用户不同编排及简单编辑后的元数据变化。尽管数据泄漏导致评估不可靠,但该问题在音乐信息检索领域尚未得到充分重视。本研究聚焦于符号音乐领域最大公开数据集之一——Lakh MIDI Dataset(LMD)的重复问题。为评估最佳去重方法,我们以LMD的Clean MIDI子集作为基准测试集,其中同一首歌的不同版本被归类在一起。我们首先测试了基于规则的方法和以往的符号音乐检索模型,并探索了使用多种增强策略的对比学习BERT模型来发现重复文件。最终提出三个不同版本的过滤列表,在最保守设置下从178,561个文件中剔除至少38,134个重复样本。

原文摘要 · Abstract (English)

A large-scale dataset is essential for training a well-generalized deep-learning model. Most such datasets are collected via scraping from various internet sources, inevitably introducing duplicated data. In the symbolic music domain, these duplicates often come from multiple user arrangements and metadata changes after simple editing. However, despite critical issues such as unreliable training evaluation from data leakage during random splitting, dataset duplication has not been extensively addressed in the MIR community. This study investigates the dataset duplication issues regarding Lakh MIDI Dataset (LMD), one of the largest publicly available sources in the symbolic music domain. To find and evaluate the best retrieval method for duplicated data, we employed the Clean MIDI subset of the LMD as a benchmark test set, in which different versions of the same songs are grouped together. We first evaluated rule-based approaches and previous symbolic music retrieval models for de-duplication and also investigated with a contrastive learning-based BERT model with various augmentations to find duplicate files. As a result, we propose three different versions of the filtered list of LMD, which filters out at least 38,134 samples in the most conservative settings among 178,561 files.

数据清洗音乐数据去重MIR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。