用深度相似性学习识别现场演唱版歌曲,准确率达87.4%。
Automatic Live Music Song Identification Using Multi-level Deep Sequence Similarity Learning
- 基于孪生卷积网络与多层级序列相似矩阵
- 在真实现场音乐数据上实现87.4%识别率
- 适合音乐信息检索与版权管理场景
本文研究自动识别现场演出歌曲的新问题:给定一段现场录音,从音乐数据库中检索出对应的录音室版本。提出一种基于相似性学习的系统,采用孪生卷积神经网络模型,通过多层级深度序列的交叉相似矩阵衡量不同音频片段间的音乐相似性。使用人工收集的定制现场音乐数据集进行测试。实验结果表明,该系统能够成功识别87.4%的现场音乐查询。
原文摘要 · Abstract (English)
This paper studies the novel problem of automatic live music song identification, where the goal is, given a live recording of a song, to retrieve the corresponding studio version of the song from a music database. We propose a system based on similarity learning and a Siamese convolutional neural network-based model. The model uses cross-similarity matrices of multi-level deep sequences to measure musical similarity between different audio tracks. A manually collected custom live music dataset is used to test the performance of the system with live music. The results of the experiments show that the system is able to identify 87.4% of the given live music queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。