统一跨模态跨语言音乐检索,用文本做桥梁打通乐谱、音频与多语种描述
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
- 通过对比学习将乐谱、音频、文本对齐到统一空间
- 在231万对音乐-文本数据上训练,跨语言泛化能力强
- 适合研究跨文化音乐检索的学者和开发者
CLaMP 3 是一个统一框架,旨在解决音乐信息检索中的跨模态与跨语言泛化难题。它采用对比学习,将乐谱、演奏信号和音频录音等主要音乐模态与多语言文本对齐至共享表示空间,实现以文本为桥梁的非对齐模态间检索。其多语言文本编码器可适应未见过的语言,具备出色的跨语言泛化能力。我们利用检索增强生成构建了 M4-RAG,一个包含 231 万组音乐-文本对的网络规模数据集,涵盖全球多样音乐传统。为推动未来研究,我们发布 WikiMT-X 基准,包含 1,000 组乐谱-音频-丰富文本描述三元组。实验表明,CLaMP 3 在多个 MIR 任务中达到当前最优性能,显著超越以往强基线,在多模态与多语言场景下表现出卓越泛化能力。
原文摘要 · Abstract (English)
CLaMP 3 is a unified framework developed to address challenges of cross-modal and cross-lingual generalization in music information retrieval. Using contrastive learning, it aligns all major music modalities--including sheet music, performance signals, and audio recordings--with multilingual text in a shared representation space, enabling retrieval across unaligned modalities with text as a bridge. It features a multilingual text encoder adaptable to unseen languages, exhibiting strong cross-lingual generalization. Leveraging retrieval-augmented generation, we curated M4-RAG, a web-scale dataset consisting of 2.31 million music-text pairs. This dataset is enriched with detailed metadata that represents a wide array of global musical traditions. To advance future research, we release WikiMT-X, a benchmark comprising 1,000 triplets of sheet music, audio, and richly varied text descriptions. Experiments show that CLaMP 3 achieves state-of-the-art performance on multiple MIR tasks, significantly surpassing previous strong baselines and demonstrating excellent generalization in multimodal and multilingual music contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。