统一音乐识别框架,让一首歌和它的版本共用同一套系统。
Unified Music Identification for Tracks and Versions

- 设计统一基准,同时评估歌曲与版本识别的准确性和抗干扰能力。
- 现有模型无法兼顾两任务,新基线模型在10秒查询下实现统一识别。
- 揭示限制识别性能的两个关键约束,为未来扩展提供方向。
给定音乐数据库,曲目识别(TI)旨在从音频片段中检索出精确匹配的曲目,而版本识别(VI)则用于检索同一首曲目的不同演绎版本。传统上,这两项任务被分别处理。然而,由于每首曲目本身可视为其最近的版本,我们探究了是否可通过版本识别系统涵盖曲目识别。这要求系统对信号处理和音频退化具有鲁棒性。为此,我们提出一个统一基准,评估各模型在两项任务上的准确性与鲁棒性。在该基准上对比七种现有模型,发现无一模型在两项任务中均兼具准确性和鲁棒性。随后,我们训练了一个针对两项任务的基线模型,证明在10秒查询条件下实现统一系统是可行的。最后,我们分析了制约模型曲目识别性能的两个关键检索约束。我们展望将此统一范式推广至其他音乐识别任务。
原文摘要 · Abstract (English)
Given a music database, track identification (TI) retrieves the exact track matching an audio excerpt, whereas version identification (VI) retrieves its musical versions. Traditionally, the two tasks have been addressed separately. However, as every track is its own closest version, we investigate whether VI can subsume TI. This requires VI systems to be robust to both signal manipulation and audio degradation. We therefore propose a unified benchmark that evaluates accuracy and robustness on each task. Comparing seven existing models on this benchmark, we show that none of them are both accurate and robust on both tasks. We then train a baseline model targeting both tasks and show that a unified system is possible with 10 s TI queries. Lastly, we characterize the two retrieval constraints that limit our model's TI performance. We envision extending this unification to other music identification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。