arXiv:2501.02511cs.CLcs.CV2025-01中稿 · NLP4MusA 2024

用封面图提取音乐情绪,构建36万条带情感描述的音乐数据集

Can Impressions of Music be Extracted from Thumbnail Images?

  • 从音乐封面图推断情绪与场景,生成非音乐属性描述
  • 构建约36万条含情境与情感的音乐描述数据
  • 新数据集提升音乐检索效果,适合内容理解研究者

近年来,基于自然语言输入的音乐检索与生成系统研究显著增多,但缺乏大规模公开的音乐数据与对应自然语言描述(音乐标注)数据集。特别是音乐播放场景、听感情绪等非音乐信息对音乐描述至关重要,却因难以直接从音频中提取而被现有数据集严重低估。为此,本文提出一种通过音乐封面图推断非音乐属性并生成音乐标注的方法,并通过人工评估验证其有效性。同时,我们构建了一个包含约36万条标注的数据集,涵盖情境与情感等非音乐信息。利用该数据集训练音乐检索模型,并在评测中证实其在音乐检索任务中的有效性。

原文摘要 · Abstract (English)

In recent years, there has been a notable increase in research on machine learning models for music retrieval and generation systems that are capable of taking natural language sentences as inputs. However, there is a scarcity of large-scale publicly available datasets, consisting of music data and their corresponding natural language descriptions known as music captions. In particular, non-musical information such as suitable situations for listening to a track and the emotions elicited upon listening is crucial for describing music. This type of information is underrepresented in existing music caption datasets due to the challenges associated with extracting it directly from music data. To address this issue, we propose a method for generating music caption data that incorporates non-musical aspects inferred from music thumbnail images, and validated the effectiveness of our approach through human evaluations. Additionally, we created a dataset with approximately 360,000 captions containing non-musical aspects. Leveraging this dataset, we trained a music retrieval model and demonstrated its effectiveness in music retrieval tasks through evaluation.

音乐理解图像生成数据集多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。