构建音乐问答新基准,让AI准确回答艺术家背景与音乐史问题
ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering
- 用320万条音乐维基文本构建向量库,支持精准检索
- 引入1000个问题的艺术家基准,使开源模型准确率提升56.8个百分点
- 适合音乐信息检索、AI音乐理解及领域适配研究者使用
大语言模型在开放域问答中表现优异,但在音乐推理方面受限于预训练数据中音乐知识稀疏。现有音乐信息检索与计算音乐学虽具备结构化与多模态理解能力,但缺乏支撑事实性与情境性音乐问答(MQA)的资源。本文提出MusWikiDB,一个包含320万条来自14.4万篇音乐相关维基页面的向量数据库,以及ArtistMus,一个涵盖500位多样化艺术家、含流派、出道年份等元数据的1000个问题基准。该资源支持对检索增强生成(RAG)在音乐问答中的系统评估。实验表明,RAG显著提升事实准确性;开源模型性能最高提升56.8个百分点(如Qwen3 8B从35.0升至91.8),接近专有模型水平。采用RAG风格微调进一步提升事实召回与上下文推理能力,在领域内与跨领域基准上均取得改进。相比通用维基语料库,MusWikiDB实现约6个百分点的准确率提升和40%的检索加速。论文发布MusWikiDB与ArtistMus,推动音乐信息检索与领域特定问答研究,为文化丰富领域中的检索增强推理奠定基础。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have transformed open-domain question answering, yet their effectiveness in music-related reasoning remains limited due to sparse music knowledge in pretraining data. While music information retrieval and computational musicology have explored structured and multimodal understanding, few resources support factual and contextual music question answering (MQA) grounded in artist metadata or historical context. We introduce MusWikiDB, a vector database of 3.2M passages from 144K music-related Wikipedia pages, and ArtistMus, a benchmark of 1,000 questions on 500 diverse artists with metadata such as genre, debut year, and topic. These resources enable systematic evaluation of retrieval-augmented generation (RAG) for MQA. Experiments show that RAG markedly improves factual accuracy; open-source models gain up to +56.8 percentage points (for example, Qwen3 8B improves from 35.0 to 91.8), approaching proprietary model performance. RAG-style fine-tuning further boosts both factual recall and contextual reasoning, improving results on both in-domain and out-of-domain benchmarks. MusWikiDB also yields approximately 6 percentage points higher accuracy and 40% faster retrieval than a general-purpose Wikipedia corpus. We release MusWikiDB and ArtistMus to advance research in music information retrieval and domain-specific question answering, establishing a foundation for retrieval-augmented reasoning in culturally rich domains such as music.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。