arXiv:2608.04543cs.SDcs.IR2026-08中稿 · the Proceedings of…

构建百万级音乐版本识别数据集,提升模型在真实场景下的鲁棒性。

Towards Robust Version Identification in the Wild: A Dataset, Benchmark, and Fine-Tuning Study

论文配图:Towards Robust Version Identification in the Wild: A Dataset, Benchmark, and Fine-Tuning Study
图 1 · 摘自论文原文
  • 构建含110万音乐版本的DiVers数据集,覆盖专业与非专业录制内容。
  • 基于DiVers训练的模型在嘈杂输入上表现显著优于传统基准。
  • 适合研究真实世界音乐识别、模型泛化能力的学者与工程师。

现有音乐版本识别(VI)数据集主要来自SecondHandSongs和Discogs等结构化元数据,以专业录音为主,与真实场景中大量业余和用户生成内容存在领域差异。为此,本文提出DiVers,一个包含超过110万音乐版本的大规模数据集,其训练-验证-测试划分与Discogs-VI-YT、SHS100K和Da-TACOS等主流数据集兼容。除标准版本级标注外,DiVers还提供自动标签(如乐器版、现场版)及段落级音频存在性预测。通过训练当前最优的VI系统评估该数据集,结果表明:在DiVers上训练的模型对声学多样性和噪声输入表现出更强鲁棒性,同时在更干净的演播室级基准上保持稳定性能。相关数据元信息、构建代码及全部实验流程均已公开,以支持可复现性。

原文摘要 · Abstract (English)

Existing datasets for musical version identification (VI) are primarily derived from curated metadata sources such as SecondHandSongs and Discogs, and are therefore dominated by professionally recorded tracks. This leads to a domain mismatch with real-world scenarios, where amateur and user-generated content is prevalent. To address this limitation, we introduce DiVers, a large-scale VI dataset comprising over 1.1 million musical versions, with train-validation-test splits compatible with established datasets such as Discogs-VI-YT, SHS100K, and Da-TACOS. In addition to standard version-level annotations, DiVers provides automatically assigned tags (e.g., instrumental, live) and segment-level predictions indicating the presence or absence of music. We evaluate the proposed dataset by training state-of-the-art VI systems. Our results show that models trained on DiVers achieve substantially improved robustness to acoustically diverse and noisy inputs, while maintaining a stable performance on cleaner, studio-quality benchmarks. We release the dataset metadata, code for its construction, and all experimental pipelines to support reproducibility.

音乐识别数据集鲁棒性音视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。