arXiv:2510.07840cs.SDeess.AS2025-10

自动构建高质量多乐器分离数据集,提升音源分离模型性能

Automatic Curation of Large-Scale, High-Quality, Multi-Category Music Source Separation Dataset

  • 通过网络爬取原始音频,用预训练模型自动筛选清洗目标乐器片段
  • 7类音源分离任务下,模型性能提升1.16dB,清洗后数据使SDR提高2.39dB
  • 适合音乐信号处理、音频生成方向研究者使用

当前多数音乐源分离(MSS)方法依赖监督学习,受限于训练数据的数量与质量。尽管网络爬取可获取大量数据,但平台级曲目标签常导致元数据不匹配,难以准确获得“音频-标签”配对。为此,我们提出ACMID:一个通过大规模原始数据爬取,并利用基于预训练音频编码器的乐器分类器进行自动清洗,从爬取曲目中过滤并聚合目标乐器干净片段,形成精炼的ACMID-Cleaned数据集。借助丰富数据,我们将传统4类分离(人声/贝斯/鼓/其他)扩展至7类(钢琴/鼓/贝斯/原声吉他/电吉他/弦乐/木管铜管),支持更高粒度的MSS系统。在先进MSS模型上的实验表明:(i) 使用ACMID-Cleaned训练的模型相比ACMID-Uncleaned,SDR性能提升2.39dB,验证了数据清洗的有效性;(ii) 将ACMID-Cleaned纳入训练,模型平均性能提升1.16dB,证明该数据集的价值。数据爬取代码、清洗模型代码及权重已公开:https://github.com/scottishfold0621/ACMID。

原文摘要 · Abstract (English)

Most current music source separation (MSS) methods rely on supervised learning, limited by training data quantity and quality. Though web-crawling can bring abundant data, platform-level track labeling often causes metadata mismatches, impeding accurate "audio-label" pair acquisition. To address this, we present ACMID: a dataset for MSS generated through web crawling of extensive raw data, followed by automatic cleaning via an instrument classifier built on a pre-trained audio encoder that filters and aggregates clean segments of target instruments from the crawled tracks, resulting in the refined ACMID-Cleaned dataset. Leveraging abundant data, we expand the conventional classification from 4-stem (Vocal/Bass/Drums/Others) to 7-stem (Piano/Drums/Bass/Acoustic Guitar/Electric Guitar/Strings/Wind-Brass), enabling high granularity MSS systems. Experiments on SOTA MSS model demonstrates two key results: (i) MSS model trained with ACMID-Cleaned achieved a 2.39dB improvement in SDR performance compared to that with ACMID-Uncleaned, demostrating the effectiveness of our data cleaning procedure; (ii) incorporating ACMID-Cleaned to training enhances MSS model's average performance by 1.16dB, confirming the value of our dataset. Our data crawling code, cleaning model code and weights are available at: https://github.com/scottishfold0621/ACMID.

音乐分离数据清洗音频处理自动生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。