构建36万首乐曲的图文数据集,补全缺失信息提升研究质量。
JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata
- 用大模型生成歌词并补全元数据,提升数据完整性。
- 通过音乐特征与元数据联合检索,准确率显著提升。
- 适合做音乐语言理解、跨模态学习的研究者使用。
我们推出了JamendoMaxCaps,一个包含超过36.2万首来自知名平台Jamendo的免费授权器乐曲的大规模音乐-文本数据集。该数据集包含由先进描述模型生成的文本描述,并通过填补缺失元数据进行了增强。我们还引入了一个检索系统,结合音乐特征和元数据识别相似歌曲,并利用本地大语言模型(LLLM)填充缺失信息。该方法有效提升了数据集的完整性和信息量,为音乐-语言理解任务提供了高质量资源。我们通过五种不同度量方式对方法进行了定量验证。公开发布后,将推动音乐检索、多模态表征学习及生成式音乐模型等方向的研究进展。
原文摘要 · Abstract (English)
We introduce JamendoMaxCaps, a large-scale music-caption dataset featuring over 362,000 freely licensed instrumental tracks from the renowned Jamendo platform. The dataset includes captions generated by a state-of-the-art captioning model, enhanced with imputed metadata. We also introduce a retrieval system that leverages both musical features and metadata to identify similar songs, which are then used to fill in missing metadata using a local large language model (LLLM). This approach allows us to provide a more comprehensive and informative dataset for researchers working on music-language understanding tasks. We validate this approach quantitatively with five different measurements. By making the JamendoMaxCaps dataset publicly available, we provide a high-quality resource to advance research in music-language understanding tasks such as music retrieval, multimodal representation learning, and generative music models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。