arXiv:2509.06936cs.SDeess.AS2025-09被引 1

用专家标注数据构建音乐自动打标新基准,对比通用标签效果。

Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets

  • 基于MGPHot专家标注数据构建标准化评测集
  • 发现专家标签与通用标签在模型表现上差异显著
  • 提供音频链接、划分方案和预计算表示,便于复现

音乐自动打标旨在为音频片段自动分配如流派、情绪或乐器等描述性标签。由于其挑战性、语义多样性及实际应用价值,已成为评估通用音乐表征性能的常见下游任务。本文引入基于最新MGPHot数据集的新基准,该数据包含音乐学专家标注,可提供额外洞察并用于与通用标签数据集结果对比。尽管原MGPHot数据已证实对计算音乐学有用,但原始数据未含音频,也无标准评测设置。为此,我们提供了可获取音频的YouTube链接,提出了标准的训练/验证/测试划分,并提供了七种先进模型的预计算表示。利用这些资源,我们在MGPHot与标准参考标签数据集上评估了多个模型,揭示了专家标签与通用标签间的显著差异。整体而言,本研究为未来音乐理解研究提供了更先进的评测框架。

原文摘要 · Abstract (English)

Music autotagging aims to automatically assign descriptive tags, such as genre, mood, or instrumentation, to audio recordings. Due to its challenges, diversity of semantic descriptions, and practical value in various applications, it has become a common downstream task for evaluating the performance of general-purpose music representations learned from audio data. We introduce a new benchmarking dataset based on the recently published MGPHot dataset, which includes expert musicological annotations, allowing for additional insights and comparisons with results obtained on common generic tag datasets. While MGPHot annotations have been shown to be useful for computational musicology, the original dataset neither includes audio nor provides evaluation setups for its use as a standardized autotagging benchmark. To address this, we provide a curated set of YouTube URLs with retrievable audio, and propose a train/val/test split for standardized evaluation, and precomputed representations for seven state-of-the-art models. Using these resources, we evaluated these models in MGPHot and standard reference tag datasets, highlighting key differences between expert and generic tag annotations. Altogether, our contributions provide a more advanced benchmarking framework for future research in music understanding.

音乐理解自动打标专家标注基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。