arXiv:2509.14891cs.MMcs.IR2025-09被引 1

构建了基于艺术家和专辑的多模态音乐数据集,支持更细粒度的音乐检索任务。

Music4All A+A: A Multimodal Dataset for Music Information Retrieval Tasks

  • 基于音乐作品集构建艺术家与专辑级多模态数据集
  • 包含6741位艺术家、19511张专辑的多模态信息,支持跨层级分析
  • 适用于音乐推荐、跨域分类等任务,尤其适合研究缺失模态问题

音乐具有音频、歌词、音乐视频等多种模态特征,推动了多模态音乐信息检索(MIR)任务的发展,如流派分类和自动标签。然而,现有数据集多聚焦于单曲级别,缺乏对艺术家和专辑层面的细粒度支持。为此,本文提出 Music4All A+A,一个基于 Music4All-Onion 数据集构建的艺术家与专辑级多模态数据集。该数据集涵盖6,741位艺术家和19,511张专辑,提供元数据、流派标签、图像表示及文本描述,并可访问原始轨道级多模态数据(包括用户-项目交互)。实验表明,图像在艺术家与专辑流派分类中更具信息量,且现有模型在跨域泛化能力上表现有限。代码已开源,数据集可通过链接获取,采用 CC BY-NC-SA 4.0 许可。

原文摘要 · Abstract (English)

Music is characterized by aspects related to different modalities, such as the audio signal, the lyrics, or the music video clips. This has motivated the development of multimodal datasets and methods for Music Information Retrieval (MIR) tasks such as genre classification or autotagging. Music can be described at different levels of granularity, for instance defining genres at the level of artists or music albums. However, most datasets for multimodal MIR neglect this aspect and provide data at the level of individual music tracks. We aim to fill this gap by providing Music4All Artist and Album (Music4All A+A), a dataset for multimodal MIR tasks based on music artists and albums. Music4All A+A is built on top of the Music4All-Onion dataset, an existing track-level dataset for MIR tasks. Music4All A+A provides metadata, genre labels, image representations, and textual descriptors for 6,741 artists and 19,511 albums. Furthermore, since Music4All A+A is built on top of Music4All-Onion, it allows access to other multimodal data at the track level, including user--item interaction data. This renders Music4All A+A suitable for a broad range of MIR tasks, including multimodal music recommendation, at several levels of granularity. To showcase the use of Music4All A+A, we carry out experiments on multimodal genre classification of artists and albums, including an analysis in missing-modality scenarios, and a quantitative comparison with genre classification in the movie domain. Our experiments show that images are more informative for classifying the genres of artists and albums, and that several multimodal models for genre classification struggle in generalizing across domains. We provide the code to reproduce our experiments at https://github.com/hcai-mms/Music4All-A-A, the dataset is linked in the repository and provided open-source under a CC BY-NC-SA 4.0 license.

多模态音乐检索数据集流派分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。