arXiv:2602.11941cs.IRcs.AI2026-02

首个细粒度音乐检索基准,含1574段授权音乐与12.5万条相关性标注。

IncompeBench: A Permissively Licensed, Fine-Grained Benchmark for Music Information Retrieval

  • 构建多阶段标注流程,确保人工一致性与数据质量。
  • 包含1574段可自由使用音乐片段、500个多样化查询及超12.5万条标注。
  • 适合评估音乐信息检索模型性能,开源可用,支持研究与产品验证。

近年来,多模态信息检索借助深度预训练模型的强大跨模态表征能力取得了显著进展。音乐信息检索(MIR)尤其在质量上大幅提升,神经音乐表征已进入日常应用。然而,高质量的音乐检索评估基准仍显不足。为此,我们提出 extbf{IncompeBench},一个精心标注的基准数据集,包含1,574段可自由使用、高质量的音乐片段,500个多样化的查询,以及超过125,000条独立的相关性判断。这些标注通过多阶段流水线生成,人评者间一致性高。数据集已公开于 https://huggingface.co/datasets/mixedbread-ai/incompebench-strict 与 https://huggingface.co/datasets/mixedbread-ai/incompebench-lenient,提示词代码见 https://github.com/mixedbread-ai/incompebench-programs。

原文摘要 · Abstract (English)

Multimodal Information Retrieval has made significant progress in recent years, leveraging the increasingly strong multimodal abilities of deep pre-trained models to represent information across modalities. Music Information Retrieval (MIR), in particular, has considerably increased in quality, with neural representations of music even making its way into everyday life products. However, there is a lack of high-quality benchmarks for evaluating music retrieval performance. To address this issue, we introduce \textbf{IncompeBench}, a carefully annotated benchmark comprising $1,574$ permissively licensed, high-quality music snippets, $500$ diverse queries, and over $125,000$ individual relevance judgements. These annotations were created through the use of a multi-stage pipeline, resulting in high agreement between human annotators and the generated data. The resulting datasets are publicly available at https://huggingface.co/datasets/mixedbread-ai/incompebench-strict and https://huggingface.co/datasets/mixedbread-ai/incompebench-lenient with the prompts available at https://github.com/mixedbread-ai/incompebench-programs.

音乐检索基准测试多模态开源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。