构建大规模音乐问答数据集,支持音乐理解与多模态研究
Jamendo-QA: A Large-Scale Music Question Answering Dataset
- 基于Jamendo平台音乐自动标注生成问答对
- 覆盖多样音乐风格与属性,支持零样本评估
- 适合音乐理解、跨模态建模与公平性研究者使用
我们提出Jamendo-QA,一个大规模音乐问答数据集。该数据集基于Jamendo平台的开源音乐作品,通过Qwen-Omni模型自动标注生成问题-答案对及与音频对齐的描述文本,支持监督训练与零样本评估。数据集涵盖多种音乐流派、乐器和元数据属性,具备良好的多样性,可支撑在不同音乐情境下的模型基准测试。我们提供了详细的数据统计,并指出潜在的流派与性别偏差,以引导公平评估。该资源具有可扩展性和公开性,旨在推动音乐理解、多模态建模及音乐导向问答系统的后续研究。
原文摘要 · Abstract (English)
We introduce Jamendo-QA, a large-scale dataset for Music Question Answering (Music-QA). The dataset is built on freely licensed tracks from the Jamendo platform and is automatically annotated using the Qwen-Omni model. Jamendo-QA provides question-answer pairs and captions aligned with music audio, enabling both supervised training and zero-shot evaluation. Our resource aims to fill the gap of music-specific QA datasets and foster further research in music understanding, retrieval, and generative applications. In addition to its scale, Jamendo-QA covers a diverse range of genres, instruments, and metadata attributes, allowing robust model benchmarking across varied musical contexts. We also provide detailed dataset statistics and highlight potential biases such as genre and gender imbalance to guide fair evaluation. We position Jamendo-QA as a scalable and publicly available benchmark that can facilitate future research in music understanding, multimodal modeling, and fair evaluation of music-oriented QA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。