arXiv:2410.13267cs.SDcs.CL2024-10NAACL被引 21

用大模型实现101种语言的跨模态音乐检索,支持文本与音频格式。

CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models

  • 基于150万组三元组预训练,融合文本、ABC记谱和MIDI多模态数据。
  • 在多语言语义搜索和跨模态分类任务上达到当前最佳性能。
  • 适合需要全球化音乐检索与多语言内容理解的研究者使用。

当前音乐信息检索系统面临语言多样性管理难与多模态融合不足的问题,限制了其在全球多模态音乐环境中的有效性。为此,我们提出CLaMP 2,支持101种语言,兼容ABC记谱(一种基于文本的乐谱格式)与MIDI(Musical Instrument Digital Interface)的音乐信息检索系统。CLaMP 2 在150万组ABC-MIDI-文本三元组上预训练,包含多语言文本编码器与多模态音乐编码器,通过对比学习进行对齐。利用大语言模型,规模化生成精细且一致的多语言描述,显著降低文本噪声并平衡语言分布。实验表明,CLaMP 2在多语言语义搜索与跨模态分类任务中均达到当前最优表现,为包容性与全球化的音乐信息检索树立新标准。

原文摘要 · Abstract (English)

Challenges in managing linguistic diversity and integrating various musical modalities are faced by current music information retrieval systems. These limitations reduce their effectiveness in a global, multimodal music environment. To address these issues, we introduce CLaMP 2, a system compatible with 101 languages that supports both ABC notation (a text-based musical notation format) and MIDI (Musical Instrument Digital Interface) for music information retrieval. CLaMP 2, pre-trained on 1.5 million ABC-MIDI-text triplets, includes a multilingual text encoder and a multimodal music encoder aligned via contrastive learning. By leveraging large language models, we obtain refined and consistent multilingual descriptions at scale, significantly reducing textual noise and balancing language distribution. Our experiments show that CLaMP 2 achieves state-of-the-art results in both multilingual semantic search and music classification across modalities, thus establishing a new standard for inclusive and global music information retrieval.

音乐检索多语言大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。