用多轨对比学习自动识别音乐采样,效果优于现有方法。
Automatic Music Sample Identification with Multi-Track Contrastive Learning
- 通过人工混音生成正样本对,设计新对比学习目标。
- 在多种音乐风格下表现优异,参考库越大效果越佳。
- 强调高质量分离音轨对任务至关重要,适合音乐分析研究者。
采样是现代音乐制作中广泛使用的技巧,即复用已有音频片段创作新内容。本文针对自动采样识别这一挑战性任务,提出一种自监督学习方法。该方法利用多轨数据集生成人工混音的正样本对,并设计新型对比学习目标。实验表明,该方法显著优于现有最优基线,在多种音乐风格下均表现稳健,且随着参考数据库中噪声歌曲数量增加,性能持续提升。此外,我们系统分析了训练流程各组件的贡献,特别指出高质量分离音轨对任务的关键作用。
原文摘要 · Abstract (English)
Sampling, the technique of reusing pieces of existing audio tracks to create new music content, is a very common practice in modern music production. In this paper, we tackle the challenging task of automatic sample identification, that is, detecting such sampled content and retrieving the material from which it originates. To do so, we adopt a self-supervised learning approach that leverages a multi-track dataset to create positive pairs of artificial mixes, and design a novel contrastive learning objective. We show that such method significantly outperforms previous state-of-the-art baselines, that is robust to various genres, and that scales well when increasing the number of noise songs in the reference database. In addition, we extensively analyze the contribution of the different components of our training pipeline and highlight, in particular, the need for high-quality separated stems for this task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。