arXiv:2502.06364cs.SDcs.LG2025-02被引 4

用人工数据训练模型,自动识别嘻哈音乐中的采样片段。

Automatic Identification of Samples in Hip-Hop Music via Multi-Loss Training and an Artificial Dataset

  • 用分离出的音源元素构建人工数据集,联合优化分类与度量学习损失。
  • 在真实采样上精度比传统方法高13%,可识别变调和变速后的样本。
  • 对半数商用歌曲能定位采样位置,误差小于5秒,适合音乐溯源研究。

采样是嘻哈与说唱等流行音乐中常见手法,即从已有作品中重用声音片段。目前已有多种服务帮助用户发现采样关联以促进音乐发现,但实现自动识别仍具挑战:采样常经变调、变速等处理,且时长可能仅数秒。进展受限于训练数据稀缺。本文提出一种卷积神经网络,基于人工生成的数据集进行训练,可有效识别商业嘻哈音乐中的真实采样。我们从非商用音乐数据库中提取人声、和声与打击乐元素,利用音频源分离技术生成带变换版本的样本,并训练模型进行指纹匹配。通过联合分类与度量学习损失优化,模型在真实采样上的精度比基于声学特征点的系统高出13%。此外,该模型能识别同时经过变调和变速处理的样本;在所测试的一半商用音乐中,其定位准确度可达5秒以内。

原文摘要 · Abstract (English)

Sampling, the practice of reusing recorded music or sounds from another source in a new work, is common in popular music genres like hip-hop and rap. Numerous services have emerged that allow users to identify connections between samples and the songs that incorporate them, with the goal of enhancing music discovery. Designing a system that can perform the same task automatically is challenging, as samples are commonly altered with audio effects like pitch- and time-stretching and may only be seconds long. Progress on this task has been minimal and is further blocked by the limited availability of training data. Here, we show that a convolutional neural network trained on an artificial dataset can identify real-world samples in commercial hip-hop music. We extract vocal, harmonic, and percussive elements from several databases of non-commercial music recordings using audio source separation, and train the model to fingerprint a subset of these elements in transformed versions of the original audio. We optimize the model using a joint classification and metric learning loss and show that it achieves 13% greater precision on real-world instances of sampling than a fingerprinting system using acoustic landmarks, and that it can recognize samples that have been both pitch shifted and time stretched. We also show that, for half of the commercial music recordings we tested, our model is capable of locating the position of a sample to within five seconds.

采样识别音频指纹深度学习嘻哈音乐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。