arXiv:2506.17815cs.SDeess.AS2025-06中稿 · ISMIR 2025被引 6

无需负样本的音乐图文预训练框架,提升模型性能与训练效率

SLAP: Siamese Language-Audio Pretraining Without Negative Samples for Music Understanding

  • 采用BYOL思想构建无负样本的音文联合预训练模型
  • 在文本-音乐检索和零样本分类上超越CLAP,下游任务表现优异
  • 支持单卡大规模训练,对批量大小不敏感,模态差距更小

联合嵌入空间通过多模态对比学习实现了音乐理解与生成的显著进展,但受限于大批次以有效利用负样本,导致内存开销巨大。此外,不同模态的嵌入分布差异显著,存在模态间隙。为此,我们提出无负样本的音文自监督预训练框架SLAP,借鉴BYOL范式,实现可扩展的多模态嵌入训练。实验表明,该模型能有效捕捉音乐与文本间的语义关联,在文本-音乐检索和零样本分类任务中优于CLAP;在多种音乐信息检索任务(如体裁、乐器分类、自动标注)中表现媲美甚至超过更大或有监督模型。同时,本方法显著降低模态间隙,提升对批量大小变化的鲁棒性,并可通过梯度累积实现单卡大规模训练。

原文摘要 · Abstract (English)

Joint embedding spaces have significantly advanced music understanding and generation by linking text and audio through multimodal contrastive learning. However, these approaches face large memory requirement limitations due to relying on large batch sizes to effectively utilize negative samples. Further, multimodal joint embedding spaces suffer from a modality gap wherein embeddings from different modalities lie in different manifolds of the embedding space. To address these challenges, we propose Siamese Language-Audio Pretraining (SLAP), a novel multimodal pretraining framework that allows learning powerful representations without negative samples. SLAP adapts the Bootstrap Your Own Latent (BYOL) paradigm for multimodal audio-text training, promoting scalability in training multimodal embedding spaces. We illustrate the ability of our model to learn meaningful relationships between music and text -- specifically, we show that SLAP outperforms CLAP on tasks such as text-music retrieval and zero-shot classification. We also observe competitive downstream performance on several MIR tasks, including with larger or supervised models (genre and instrument classification, auto-tagging). Additionally, our approach has attractive properties, such as a quantifiably reduced modality gap and improved robustness to batch size variations on retrieval performance. Finally, its novel formulation unlocks large-scale training on a single GPU through gradient accumulation.

音文预训练对比学习多模态单卡训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。