构建3.2万组音乐语义描述数据,更贴近真实听感表达。
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
- 从Reddit自然讨论中收集音乐-语言配对,覆盖多元语义
- 包含32,493组数据,涵盖五类音乐描述维度
- 推动人本化音乐多模态模型研究,适合生成与检索任务
音乐表征学习是音乐信息检索与生成的核心。尽管多模态学习的进展提升了文本与音频在跨模态音乐检索、文本到音乐生成及音乐到文本生成等任务中的对齐能力,现有模型仍难以捕捉用户在自然语言中表达的音乐意图。这表明当前训练与评估所用数据集未能充分反映人类描述音乐时的广泛且自然的交流方式。本文提出MusicSem,一个包含32,493组语言-音频配对的数据集,源自社交媒体平台Reddit上的有机音乐相关讨论。相比现有数据集,MusicSem涵盖了更广泛的音乐语义,反映了听众以细腻且以人为本的方式描述音乐的真实语境。为系统化这些表达,我们提出五个语义类别:描述性、氛围性、情境性、元数据相关性与上下文性。除数据集构建、分析与发布外,我们利用MusicSem评估多种多模态模型在检索与生成任务中的表现,凸显建模细粒度语义的重要性。总体而言,MusicSem为未来面向人类对齐的多模态音乐表征学习研究提供了新的语义感知资源。
原文摘要 · Abstract (English)
Music representation learning is central to music information retrieval and generation. While recent advances in multimodal learning have improved alignment between text and audio for tasks such as cross-modal music retrieval, text-to-music generation, and music-to-text generation, existing models often struggle to capture users' expressed intent in natural language descriptions of music. This observation suggests that the datasets used to train and evaluate these models do not fully reflect the broader and more natural forms of human discourse through which music is described. In this paper, we introduce MusicSem, a dataset of 32,493 language-audio pairs derived from organic music-related discussions on the social media platform Reddit. Compared to existing datasets, MusicSem captures a broader spectrum of musical semantics, reflecting how listeners naturally describe music in nuanced and human-centered ways. To structure these expressions, we propose a taxonomy of five semantic categories: descriptive, atmospheric, situational, metadata-related, and contextual. In addition to the construction, analysis, and release of MusicSem, we use the dataset to evaluate a wide range of multimodal models for retrieval and generation, highlighting the importance of modeling fine-grained semantics. Overall, MusicSem serves as a novel semantics-aware resource to support future research on human-aligned multimodal music representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。